1
00:00:00,530 --> 00:00:09,200
Hi, in this video, I'm going to show you how to create a plot of daily average ratings so our graph

2
00:00:09,200 --> 00:00:13,800
will show the average rating for each day. To do that,

3
00:00:13,850 --> 00:00:20,060
I am not going to work on the previous Jupyter file because I like to keep things separate.

4
00:00:20,090 --> 00:00:23,530
So for the visualization part, I'm going to create a new notebook.

5
00:00:23,540 --> 00:00:24,980
I would suggest you do the same.

6
00:00:27,380 --> 00:00:29,090
So first of all, we need to load

7
00:00:30,430 --> 00:00:36,220
the data frame, so I'm just going to copy the cell from the previous notebook.

8
00:00:38,520 --> 00:00:46,450
And paste it there and just print quickly the heads of the data just to double-check everything

9
00:00:46,470 --> 00:00:47,070
is all right.

10
00:00:47,460 --> 00:00:47,820
Yes.

11
00:00:47,910 --> 00:00:53,610
So these are the data, the raw data. Then press escape and press be on your keyboard to create a

12
00:00:53,610 --> 00:00:55,950
new cell, press enter to enter the new cell.

13
00:00:55,950 --> 00:00:57,870
And let's create that graph.

14
00:00:58,170 --> 00:01:04,650
But before we create the graph, we need to do some sort of data aggregation.

15
00:01:04,990 --> 00:01:05,320
Right.

16
00:01:05,320 --> 00:01:09,900
So we are turning data into information, but our data are quite raw.

17
00:01:10,080 --> 00:01:19,800
You see that for one day we have multiple ratings left by different students, for example, on the second

18
00:01:19,800 --> 00:01:21,230
of April.

19
00:01:21,240 --> 00:01:23,040
So second of April,

20
00:01:23,280 --> 00:01:24,390
we have this rating.

21
00:01:24,390 --> 00:01:26,810
We have that and that and that.

22
00:01:26,940 --> 00:01:32,700
And so on. We want to aggregate those numbers into one average rating.

23
00:01:32,880 --> 00:01:33,930
How do we do that?

24
00:01:34,500 --> 00:01:35,400
We do that.

25
00:01:35,400 --> 00:01:43,460
We can do that by using the Pandas groupby method which produces a new data frame.

26
00:01:43,470 --> 00:01:45,300
So it's going to be a new data frame.

27
00:01:45,570 --> 00:01:47,700
But we've aggregated data.

28
00:01:47,700 --> 00:01:52,080
So some new data which are which are averages of the raw data frame.

29
00:01:53,020 --> 00:01:55,160
Let me show you how groupby works.

30
00:01:55,620 --> 00:02:00,530
So I said that groupby produces a new data frame.

31
00:02:00,540 --> 00:02:05,700
Therefore, I'm going to create the new variable where the new data frame is going to be stored.

32
00:02:05,880 --> 00:02:09,610
So day_average is my new variable and that would be equal to data.

33
00:02:10,200 --> 00:02:18,000
So that data frame, groupby.
I really like this method, it's very intuitive.

34
00:02:18,010 --> 00:02:25,820
So you are grouping these data by... What do we want to group by?

35
00:02:26,430 --> 00:02:28,470
Well, by timestamp.

36
00:02:31,310 --> 00:02:33,270
Let's see what that gives us.

37
00:02:33,740 --> 00:02:36,290
So I'm going to print out the average.

38
00:02:37,740 --> 00:02:44,480
In here, the head of the average executes and let's see what we got.

39
00:02:47,870 --> 00:02:57,230
We got basically the same data frame that is because groupby was not able to group the data by timestamp

40
00:02:57,440 --> 00:03:06,710
because the way groupby works is that groupby tries to find identical values in that given column.

41
00:03:07,490 --> 00:03:13,800
In this case, timestamped doesn't have any identical values because each value is unique.

42
00:03:13,820 --> 00:03:20,270
You can see it's the same Debord's this review was left at this time.

43
00:03:21,010 --> 00:03:22,820
The other review was left at this time.

44
00:03:22,970 --> 00:03:25,180
And so on. So each value is different.

45
00:03:25,460 --> 00:03:30,410
Therefore, before we apply group by, we need to do some processing here.

46
00:03:31,340 --> 00:03:35,950
We need to add a new column in the data data frame.

47
00:03:36,410 --> 00:03:44,210
I'm going to name this column 'Day' with Capital D. I'm just trying to be consistent with the names

48
00:03:44,210 --> 00:03:47,810
of the columns, since these are with capital letters.

49
00:03:47,840 --> 00:03:49,010
They start with a capital letter.

50
00:03:49,010 --> 00:03:56,690
I'm going to create this new column with a capital letter so data Day equal to data

51
00:03:58,600 --> 00:03:59,800
Timestamp

52
00:04:01,870 --> 00:04:13,300
dot dt, dt is a property which gives us access to a number of data time attributes such as date.

53
00:04:14,890 --> 00:04:20,330
Sorry date, you can do month and so on.

54
00:04:20,350 --> 00:04:21,860
For now, we need the date.

55
00:04:23,110 --> 00:04:24,990
Let me comment this out.

56
00:04:24,990 --> 00:04:31,330
So I'm going to select them and press control slash or command slash to comment them out.

57
00:04:31,750 --> 00:04:36,810
And I'm going to show you the new data frame.

58
00:04:37,450 --> 00:04:39,580
So that is the data frame.

59
00:04:41,600 --> 00:04:48,080
So what I just did is that I extracted from this timestamp, I extracted the date only.

60
00:04:49,040 --> 00:04:53,080
Therefore what I got is that for each timestamp I got the date.

61
00:04:53,360 --> 00:04:58,650
So for this timestamp is 2nd of April, 2nd of April and so on.

62
00:04:58,790 --> 00:05:01,510
So this way we got some identical data.

63
00:05:02,540 --> 00:05:08,700
You know, if you wanted a month, you'd get the number of the month.

64
00:05:09,290 --> 00:05:12,470
So 4, 4, 4 and so one.

65
00:05:13,190 --> 00:05:14,470
But we need date.

66
00:05:14,480 --> 00:05:16,180
So I'm going to keep the date there.

67
00:05:16,860 --> 00:05:25,150
Now, we can uncommentent these and let's delete the data dot head because we don't need it anymore.

68
00:05:25,850 --> 00:05:30,050
And now we can try this groupby method again.

69
00:05:30,050 --> 00:05:32,030
But be careful this time.

70
00:05:32,040 --> 00:05:33,190
We need day here.

71
00:05:33,200 --> 00:05:35,870
As we said, we want to groupby day.

72
00:05:36,530 --> 00:05:38,120
Let's execute and see what we get.

73
00:05:42,610 --> 00:05:45,710
Maybe it's not the result we were waiting for.

74
00:05:45,730 --> 00:05:52,550
So you see that day is not yet aggregated because we need to give another command here.

75
00:05:52,720 --> 00:05:57,580
We need to tell pandas the methods of aggregation.

76
00:05:57,610 --> 00:06:06,610
So do you want to aggregate based on the mean or the count? In this case is the mean.

77
00:06:06,630 --> 00:06:11,680
So we'd say .mean(), the method, execute.

78
00:06:11,890 --> 00:06:14,680
And to this time this is what we got.

79
00:06:16,670 --> 00:06:22,820
So that's just the head, but if you print out the entire data frame, you see that it goes like

80
00:06:22,820 --> 00:06:25,330
that up to this date.

81
00:06:26,210 --> 00:06:33,270
So for each row, we have one day, 1st of January, 2nd of January, 3rd of January, and so on.

82
00:06:33,410 --> 00:06:38,690
So that is the average rating of all the causes for that day.

83
00:06:40,870 --> 00:06:44,480
No, you need to understand this product.

84
00:06:44,500 --> 00:06:48,910
This is, as I told you, it's a data frame type.

85
00:06:50,530 --> 00:06:53,200
So Pandas data frame,

86
00:06:55,970 --> 00:06:58,460
but this has one column.

87
00:07:00,020 --> 00:07:07,730
So that is not a column, day is not a column, day is actually the index, you see that if you say

88
00:07:07,730 --> 00:07:08,810
day average dot columns

89
00:07:11,300 --> 00:07:18,510
rating is the only column, and if you want to access rating, you do like that, right?

90
00:07:18,520 --> 00:07:19,830
And you get this series.

91
00:07:21,150 --> 00:07:22,070
Now, what is this?

92
00:07:22,120 --> 00:07:23,290
This is the index.

93
00:07:24,430 --> 00:07:30,550
Therefore, if you want to access the daily column, you don't do it like that because that is a syntax

94
00:07:30,550 --> 00:07:31,750
to access columns.

95
00:07:31,990 --> 00:07:38,560
When you want to access that special column, so that index column, you want to see that index.

96
00:07:40,610 --> 00:07:45,160
And then we get this series which is

97
00:07:48,670 --> 00:07:56,550
type index, but you can easily convert it into a list, for example, if you like to plot them.

98
00:07:57,790 --> 00:08:03,460
So just like columns, indexes such as this one are also list like types.

99
00:08:04,150 --> 00:08:05,360
So arrays of data.

100
00:08:06,220 --> 00:08:10,410
Let's now do the plotting. To do the plotting

101
00:08:10,420 --> 00:08:18,250
we are going to need the Matplotlib library, so I am going to impoirt it in here.

102
00:08:19,870 --> 00:08:25,090
So import matplotlib.pyplot.

103
00:08:25,810 --> 00:08:35,530
We need that module of the library and a good practice is to as plt so you will see on the web that

104
00:08:35,530 --> 00:08:36,670
everyone uses plt.

105
00:08:36,820 --> 00:08:43,180
So if you want to be consistent with other programmers, you want to import that as plt.

106
00:08:43,450 --> 00:08:48,370
That also makes your job easier because you don't have to type that down.

107
00:08:48,580 --> 00:08:52,390
But you can just say plt as we will do here.

108
00:08:52,690 --> 00:09:03,990
So plt.plot is the method and this method basically gets two arguments, the X and the Y.

109
00:09:04,150 --> 00:09:10,540
So we are building a graph with an X and Y axis. Along the X axis

110
00:09:10,540 --> 00:09:13,090
we are going to have the days.

111
00:09:13,400 --> 00:09:18,640
That means we want the average.index.

112
00:09:19,120 --> 00:09:20,620
So that array there.

113
00:09:23,370 --> 00:09:31,410
And along Y we want day_average rating column.

114
00:09:32,970 --> 00:09:33,570
Execute.

115
00:09:36,650 --> 00:09:44,610
Oh, I got an error.  Plt is not defined because I forgot to execute this cell so that the import is valid.

116
00:09:44,810 --> 00:09:48,500
Now I can execute this again and this is the product.

117
00:09:49,280 --> 00:09:56,390
So along the Y axis, we have dates, which I know they are a bit unvisible, but we are going to

118
00:09:56,390 --> 00:09:56,960
fix that.

119
00:09:57,770 --> 00:10:01,850
And along the Y axis, we have the ratings column.

120
00:10:02,150 --> 00:10:09,590
And so you see that this starts from three point eight somewhere here, up to five now.

121
00:10:09,590 --> 00:10:14,780
Matplotlib picks that range automatically by looking at the data.

122
00:10:15,380 --> 00:10:23,990
So our rating column if you take a look and it on it, you say day_average rating.

123
00:10:24,380 --> 00:10:30,620
If you extend the max, the maximum value, you'll see that it is 5.0.

124
00:10:31,100 --> 00:10:35,840
And if you see the minimum, you see that it's around three point eight.

125
00:10:35,870 --> 00:10:42,400
So one day students left a 5.0 average rating.

126
00:10:42,420 --> 00:10:45,740
So all of them and another day they left three point eight.

127
00:10:45,750 --> 00:10:54,590
So matplotlib is putting those two as the limits of the Y axis instead of starting the axis from zero

128
00:10:54,740 --> 00:10:55,760
up to five.

129
00:10:55,940 --> 00:11:00,360
And that would make the plot less readable.

130
00:11:00,860 --> 00:11:03,610
So that's so that's a good thing of matplotlib.

131
00:11:04,070 --> 00:11:08,720
The bad thing is, as you can see, these graph is not interactive.

132
00:11:08,720 --> 00:11:11,090
So it's just an image file.

133
00:11:11,090 --> 00:11:14,360
You can not have pop up capabilities.

134
00:11:14,360 --> 00:11:19,180
So you could see some values if you if you over your or your mouse somewhere.

135
00:11:19,580 --> 00:11:21,620
So matplotlib cannot do that.

136
00:11:22,200 --> 00:11:35,930
However, we can improve this a little bit by declaring a figure object and give it a fixed size argument

137
00:11:35,930 --> 00:11:39,410
of let's say twenty five, three.

138
00:11:39,830 --> 00:11:45,780
That is the width and that is the height of the plot.

139
00:11:46,520 --> 00:11:53,100
So now you can see that we have a longer X axis and the shorter Y axis.

140
00:11:53,990 --> 00:12:00,440
Now if you don't agree with this graph, if you think that this is still not readable, it doesn't tell

141
00:12:00,440 --> 00:12:07,790
you much about the trend, so if the rating has been increasing with time or not.

142
00:12:09,720 --> 00:12:17,910
Then what we could do is we could downsampled the data, so instead of extracting daily averages, we

143
00:12:17,910 --> 00:12:26,590
could extract weekly averages and therefore we would have less points along the x axis and the smoother

144
00:12:26,610 --> 00:12:27,120
line.

145
00:12:28,660 --> 00:12:33,220
So in my opinion, these are too much data, it's unreadable, it's not useful.

146
00:12:33,750 --> 00:12:37,400
Let's down sample them with a better graph in the next video.

147
00:12:38,530 --> 00:12:41,250
So this is what you learned in this video.

148
00:12:41,260 --> 00:12:44,260
You learned how to group data and you learned how to plot them.

149
00:12:44,590 --> 00:12:51,280
Now, let me make a small revision of the groupby methods in case you are still confused.

150
00:12:51,460 --> 00:12:57,640
So let me remove that in another cell here and let me print out

151
00:13:00,180 --> 00:13:07,080
the head of the aggregated data frame, so you see we have rating and we have the index here and what

152
00:13:07,080 --> 00:13:13,010
happened with course name with this column, what happened with the comments column?

153
00:13:13,020 --> 00:13:15,010
What happened with Timestamp column?

154
00:13:15,750 --> 00:13:26,220
Well, they disappeared because this mean method only works with columns that have number values, such

155
00:13:26,220 --> 00:13:27,210
as the ratings.

156
00:13:27,210 --> 00:13:34,020
So it cannot calculate an average on that column comment or course name, or timestamp.

157
00:13:35,250 --> 00:13:42,240
Therefore, columns such as comment, timestamp and course name will be ultimately dropped, deleted

158
00:13:42,240 --> 00:13:47,450
by this method, and only those columns will be kept.

159
00:13:47,550 --> 00:13:52,050
Similarly you could do a count instead of that.

160
00:13:52,350 --> 00:13:57,600
In that case, you'd get a different data frame.

161
00:13:57,600 --> 00:14:04,230
So you see that we have the count of course names, which means that 46

162
00:14:05,980 --> 00:14:09,200
rows have been created for that date.

163
00:14:09,460 --> 00:14:16,930
In other words, 46 reviews, so we have 46 the ratings, whatever you like to call it, we have seven

164
00:14:16,930 --> 00:14:24,010
here because what Pandas does is that it doesn't take into account NaN values such as this one and that

165
00:14:24,010 --> 00:14:24,150
one.

166
00:14:24,160 --> 00:14:30,120
And so one, therefore, we could have this count plot.

167
00:14:30,610 --> 00:14:35,410
So we show how many ratings we had each day.

168
00:14:37,760 --> 00:14:41,390
And that is what I wanted to teach you in this video.

169
00:14:41,420 --> 00:14:42,020
Thanks a lot.

170
00:14:42,050 --> 00:14:42,850
I'll talk to you later.

