1
00:00:01,330 --> 00:00:02,230
Hi, welcome back.

2
00:00:02,260 --> 00:00:09,010
In this video, we are going to generate a plot containing different lines and each line is going to

3
00:00:09,010 --> 00:00:14,030
represent the average rating by month for a particular course.

4
00:00:14,260 --> 00:00:15,980
So we have multiple courses.

5
00:00:16,000 --> 00:00:18,160
We are going to have one line for each course.

6
00:00:18,800 --> 00:00:19,540
Let's do that.

7
00:00:20,020 --> 00:00:25,280
Let's put a heading h3.

8
00:00:25,540 --> 00:00:31,150
This would be average rating by month by course.

9
00:00:32,630 --> 00:00:32,990
Right.

10
00:00:33,950 --> 00:00:35,400
This is how we do that.

11
00:00:35,660 --> 00:00:36,920
First of all, we need

12
00:00:39,590 --> 00:00:43,970
a month column in the data data frame.

13
00:00:44,150 --> 00:00:46,150
Now, I know that we did that here.

14
00:00:46,160 --> 00:00:51,920
We already have that month column, but I want to keep the code separate so it doesn't harm if you do

15
00:00:51,920 --> 00:00:53,860
it again here just for consistency.

16
00:00:54,140 --> 00:00:58,810
So to see that this cell is doing that particular thing.

17
00:00:59,000 --> 00:01:01,640
So it's building a graph of multiple lines.

18
00:01:01,820 --> 00:01:04,080
And for that graph, we need to this expression.

19
00:01:04,080 --> 00:01:05,630
So it's the same expression.

20
00:01:05,810 --> 00:01:07,940
I can just copy and paste it in here.

21
00:01:09,140 --> 00:01:09,490
Right.

22
00:01:09,500 --> 00:01:11,210
And we don't need to change anything there.

23
00:01:11,210 --> 00:01:13,190
So we just need to extract the month

24
00:01:15,800 --> 00:01:17,290
and let's check.

25
00:01:17,960 --> 00:01:20,480
So we have the month for each row.

26
00:01:21,350 --> 00:01:24,500
Since we are grouping by months, it makes sense to have that.

27
00:01:25,730 --> 00:01:31,730
Then we need to use the groupby method to create a new data frame, which would be an aggregation.

28
00:01:32,090 --> 00:01:42,260
Now, just so that we compare it to compare it to let's see what the previous month average data frame

29
00:01:42,260 --> 00:01:43,160
looks like.

30
00:01:43,700 --> 00:01:47,300
So again, that is what we had.

31
00:01:47,990 --> 00:01:52,180
This data frame is missing the information about the courses.

32
00:01:52,400 --> 00:01:54,860
So it's missing the course name information.

33
00:01:55,130 --> 00:02:02,570
You see here we have this column that we need to have that dimension somehow in our new aggregated data

34
00:02:02,570 --> 00:02:02,930
frame.

35
00:02:03,620 --> 00:02:06,370
So what we need to do here is something different.

36
00:02:07,040 --> 00:02:11,090
Again, we need to create a new variable.

37
00:02:11,270 --> 00:02:16,820
Let's call this a month_average_crs maybe for course.

38
00:02:17,030 --> 00:02:25,040
So to distinguish between that data frame and that month_average, month_average_crs.

39
00:02:25,040 --> 00:02:25,820
This again will be

40
00:02:27,750 --> 00:02:31,860
a product, a byproduct of that data frame.

41
00:02:32,490 --> 00:02:39,810
So using the groupby methods, but in this case, let me delete that so you can see the other codes here

42
00:02:40,080 --> 00:02:42,800
to see the difference. In this case,

43
00:02:44,400 --> 00:02:52,830
we are going to have not only the month here in the list of columns to be used for the aggregation process,

44
00:02:53,280 --> 00:02:57,480
but we are also going to have can you guess what? Course name.

45
00:02:58,530 --> 00:03:02,330
So by month, by course. That's what we need.

46
00:03:03,600 --> 00:03:07,950
Let's apply now mean() and and see what we get so far.

47
00:03:14,190 --> 00:03:24,180
So this is a data frame and basically this one has actually two levels of indexes and it has the month

48
00:03:24,180 --> 00:03:28,830
and the course name, you can see that if you do that index.

49
00:03:30,270 --> 00:03:31,980
So you see it's a multi index.

50
00:03:32,820 --> 00:03:34,620
You see the month column.

51
00:03:35,910 --> 00:03:42,420
And also the course name, which doesn't show here, but it's there and as you saw it, you could

52
00:03:42,420 --> 00:03:43,620
see it in here.

53
00:03:43,910 --> 00:03:45,960
You see that these are in bold.

54
00:03:47,640 --> 00:03:49,990
These values, that that means they are indexes.

55
00:03:51,120 --> 00:03:54,450
This here is a column, so the columns are

56
00:03:57,540 --> 00:03:59,820
rating only, only one column.

57
00:04:01,080 --> 00:04:10,620
Now, this is not very useful yet because of the way this is constructed is that we have a row here,

58
00:04:10,950 --> 00:04:13,130
so this is a group of rows.

59
00:04:13,620 --> 00:04:14,720
It starts here.

60
00:04:14,790 --> 00:04:19,140
So the first course, is the second course, the third, fourth, the fifth.

61
00:04:19,500 --> 00:04:25,980
And there's also some other courses here, which actually we can see them by applying here as lines.

62
00:04:26,170 --> 00:04:28,770
So let's see the first 20 records.

63
00:04:31,560 --> 00:04:41,730
Now we see the first 20 records in full, so complete first day, sorry first month of 2018, second

64
00:04:41,730 --> 00:04:51,510
month of 2018 and so on, so that course has that rating for that month, that course has that average

65
00:04:51,510 --> 00:04:53,360
rating for that month.

66
00:04:53,880 --> 00:05:01,800
And so on, all the other courses. Then the same pattern such as this one here is repeated over and over

67
00:05:01,800 --> 00:05:10,070
again for all the months. Now to have this data are in a better structure,

68
00:05:10,500 --> 00:05:20,940
what we would do is we would apply here and dot unstack method to basically unstack this data frame

69
00:05:21,510 --> 00:05:24,450
and end up with this better structure.

70
00:05:25,410 --> 00:05:33,180
So now we have sort of a pivot table, so to say. 
We have the month here, first month, second month,

71
00:05:33,180 --> 00:05:37,710
third month, and each column now is representing a course.

72
00:05:41,020 --> 00:05:47,920
So if you want to know, for example, the rating of 'The Complete Python Course: Build 10 Profesional

73
00:05:47,920 --> 00:05:49,390
OOP Apps for

74
00:05:50,420 --> 00:05:58,700
a particular month, let's say this one, it's NaN because the cause was not published yet in that date,

75
00:05:58,940 --> 00:06:04,730
but if you look at the end, let's say minus 20,

76
00:06:11,440 --> 00:06:19,450
then you see that from the first month of 2020, we have a rating for that course, an average

77
00:06:19,450 --> 00:06:20,760
rating, right?

78
00:06:24,550 --> 00:06:27,860
And so on, what can we do this data frame now?

79
00:06:28,570 --> 00:06:36,280
Well, we can plot it, but this time we are going to use a different approach because we cannot use

80
00:06:36,280 --> 00:06:39,310
that plt.plot.

81
00:06:41,330 --> 00:06:49,100
Since we have multiple, multiple columns, so this expects an X and a Y, but which one is the X

82
00:06:49,640 --> 00:06:49,890
right?

83
00:06:49,920 --> 00:06:54,020
The X of course is month, we are clear about that.

84
00:06:54,890 --> 00:06:59,830
But which one is the Y? Is it that or that or that?

85
00:07:00,080 --> 00:07:02,270
So we have multiple columns.

86
00:07:02,450 --> 00:07:09,860
One way would be to write this plot function multiple times or maybe create a loop that iterate through

87
00:07:09,860 --> 00:07:10,550
the data frame.

88
00:07:10,550 --> 00:07:18,620
But another easier way is to just point to average_crs.plot.

89
00:07:20,380 --> 00:07:28,740
And voila, we have of course, it looks a mess, but let's improve it a bit since we used the plot

90
00:07:28,750 --> 00:07:38,110
function directly from the data frame, not from the plt, then we can use a figsize argument in here.

91
00:07:38,410 --> 00:07:40,720
So let's say 25, 3.

92
00:07:42,360 --> 00:07:44,280
And now it looks a bit better.

93
00:07:46,220 --> 00:07:47,550
Still not ideal.

94
00:07:48,050 --> 00:07:50,630
Maybe we could increase that to eight.

95
00:07:52,740 --> 00:08:00,270
OK, now it's working so that we have a legend that shows the color of each of the causes.

96
00:08:01,510 --> 00:08:08,220
Again, I'm not a fan of this kind of graph, but when we use the Highcharts library with the Web interface,

97
00:08:08,320 --> 00:08:10,390
this is going to look much better.

98
00:08:10,840 --> 00:08:13,060
Now, what if we use count here?

99
00:08:17,460 --> 00:08:25,560
We would get the graph correctly, but the legend would be a bit of a mess, that is because in the

100
00:08:28,830 --> 00:08:35,130
month_average_crs dataframe

101
00:08:40,480 --> 00:08:49,540
we have not only the counts of the ratings, but we also have other kinds of timestamps and things

102
00:08:49,540 --> 00:08:50,460
we don't need.

103
00:08:50,770 --> 00:08:54,360
You see that, for example '100 Exercises I'

104
00:08:54,370 --> 00:08:57,400
has the time stamp here.

105
00:08:57,400 --> 00:09:00,330
You see it's on another level of columns.

106
00:09:00,610 --> 00:09:04,570
So all these belong to timestamp, right?

107
00:09:05,050 --> 00:09:07,380
The count of time stamp.

108
00:09:08,890 --> 00:09:13,240
And then we have the count of rating starting from here and up to somewhere.

109
00:09:13,360 --> 00:09:17,500
We don't see all the columns because Jupiter is truncating them.

110
00:09:18,310 --> 00:09:20,840
But you get the idea day and week.

111
00:09:21,640 --> 00:09:24,280
So how do we extract only the rating?

112
00:09:24,430 --> 00:09:25,950
Well, that is easy, actually.

113
00:09:26,830 --> 00:09:39,700
I told you that groupby which starts in here and ends in here, groupby returns a data frame, now out

114
00:09:39,700 --> 00:09:40,660
of a data frame

115
00:09:40,660 --> 00:09:43,300
we extract only the rating.

116
00:09:46,860 --> 00:09:56,130
And problem solved now we get a clear graph showing the number of ratings left for each month of the

117
00:09:56,130 --> 00:09:56,460
year.

118
00:09:58,200 --> 00:10:05,760
So starting from the very beginning, 2018, up to 2020.
You see this pink line here, which is the

119
00:10:05,760 --> 00:10:14,690
new course, 'The Complete Python Course', it starts somewhere in 2021, so the first month of 2021.

120
00:10:15,360 --> 00:10:22,770
And of course, you can again get the mean, just like you did previously, but with mean we again get

121
00:10:22,770 --> 00:10:30,500
the same results since the mean method ignores those timestamps of those nonnumeric columns.

122
00:10:30,720 --> 00:10:35,850
So it's any way it drops them, but count does not drop them.

123
00:10:35,850 --> 00:10:39,680
It's counts them up and gives us so much data.

124
00:10:40,380 --> 00:10:40,700
Right.

125
00:10:40,800 --> 00:10:41,820
I hope this was clear.

126
00:10:41,820 --> 00:10:43,300
And I'll talk to you later.

127
00:10:43,350 --> 00:10:43,710
See you.

