1
00:00:00,590 --> 00:00:07,250
Hi, welcome back to the previous video, we developed a plot showing daily averages of ratings,

2
00:00:07,250 --> 00:00:10,970
but also the daily number of ratings. In this video,

3
00:00:10,980 --> 00:00:14,640
we want to rebuild this graph but on a weekly basis.

4
00:00:14,840 --> 00:00:15,520
Let's do that.

5
00:00:15,800 --> 00:00:17,690
So let's organize this data frame a bit.

6
00:00:18,830 --> 00:00:21,290
I'm going to click here.

7
00:00:21,870 --> 00:00:34,940
Press enter, press m, enter and give this section a title.
Rating average by day slash count since we did both

8
00:00:34,940 --> 00:00:37,280
averages and counts. Execute that cell.

9
00:00:37,290 --> 00:00:45,320
Now we need a new cell here, B, rating average by week.

10
00:00:47,180 --> 00:00:49,880
Again, make sure you have access to the data variable.

11
00:00:50,030 --> 00:00:52,520
So this is my data data frame.

12
00:00:52,790 --> 00:00:58,570
Otherwise you have to run this cell first and then perform the next operations.

13
00:00:58,880 --> 00:01:06,020
So first thing we want to do is we want to I'm going to copy that line.

14
00:01:06,650 --> 00:01:12,410
And so what we want to change from this is first we need to give to create a new 

15
00:01:14,670 --> 00:01:23,060
column in the data data frame, so let's name it a week, and we extract from timestamp.dt.

16
00:01:23,720 --> 00:01:26,050
Should it be week?

17
00:01:26,790 --> 00:01:28,110
Let's try that out.

18
00:01:30,410 --> 00:01:32,270
Or just print out the entire data frame.

19
00:01:34,060 --> 00:01:40,570
All right, we get a warning, let's ignore that for a while and we get the week column.

20
00:01:41,710 --> 00:01:51,220
So we see this is the 13th week, that's the first week, oh, let's see, actually what is the maximum

21
00:01:53,440 --> 00:01:54,630
of that column?

22
00:01:54,910 --> 00:01:56,500
So it's 53.

23
00:01:56,500 --> 00:01:57,560
What is the minimum?

24
00:01:57,580 --> 00:01:58,900
It's one, I suppose.

25
00:01:59,540 --> 00:01:59,770
Hmm.

26
00:02:01,270 --> 00:02:11,590
So that means we have only three weeks, so what PANDAS is doing is that it's aggregating the weeks

27
00:02:12,040 --> 00:02:19,720
of different years, which means that it is aggregating the ratings of, let's say, the first week

28
00:02:19,720 --> 00:02:28,340
of 2018 with the ratings of the first week of 2019, the ratings of the first week of to 2020.

29
00:02:28,900 --> 00:02:35,050
So all those three weeks are going to be aggregated into one and that is going to be week one.

30
00:02:35,440 --> 00:02:41,770
And then we get week two for the second week of 2018, 2019 and 2020 and so on.

31
00:02:41,980 --> 00:02:50,440
So in total, we had fifty three weeks because that's how many weeks a year has.

32
00:02:50,450 --> 00:02:57,680
Well, normally it has 52 weeks as far as I know, but maybe one of these years it's probably a leap year.

33
00:02:57,820 --> 00:03:00,690
So it has fifty three weeks I suppose.

34
00:03:00,940 --> 00:03:02,360
Anyway you get the idea.

35
00:03:03,130 --> 00:03:06,190
And so this is not what we need.

36
00:03:06,490 --> 00:03:09,400
So let's try what the warning is suggesting.

37
00:03:09,520 --> 00:03:12,670
Dt dot is, so calendar

38
00:03:15,020 --> 00:03:16,100
dot week.

39
00:03:18,350 --> 00:03:21,790
Hmmm, it seems the maximum again is 53.

40
00:03:23,570 --> 00:03:27,560
So still we are getting aggregated weeks.

41
00:03:27,920 --> 00:03:42,050
Therefore the solution here is to use strftime, which means string from time, and that gets as argument

42
00:03:43,900 --> 00:03:54,490
some date time code, such as percentage Y, maybe a dash or that, or space that is up to you, that

43
00:03:54,490 --> 00:03:55,330
is optional.

44
00:03:55,540 --> 00:04:04,900
But this is the code that you have to use if you want to extract the year and then we extract the week number.

45
00:04:04,900 --> 00:04:08,230
Let's see what we get this time.

46
00:04:09,790 --> 00:04:11,110
Now we are talking.

47
00:04:11,710 --> 00:04:15,970
So this time we get the year and we give the week.

48
00:04:17,470 --> 00:04:22,450
Therefore we can distinguish between different rows.

49
00:04:22,450 --> 00:04:26,800
So week 13 of 2021 is now under this name.

50
00:04:26,800 --> 00:04:28,270
It's not just 13.

51
00:04:29,050 --> 00:04:35,030
Therefore we would have again 2020-13, 2019-13 and so on.

52
00:04:35,140 --> 00:04:43,410
So we have unique names for weeks now and that is what helped us to construct this date time format.

53
00:04:43,720 --> 00:04:52,150
You could use other codes, so if you wanted the month, you'd use a lower M and then we would get the

54
00:04:52,150 --> 00:04:53,150
number of the month.

55
00:04:53,150 --> 00:04:59,800
So April here and the week, the number of the week, as I told you, this is optional.

56
00:04:59,960 --> 00:05:08,230
So instead of the dash, you could use a space and you get the space between the month and the week.

57
00:05:08,710 --> 00:05:10,510
Where can you find this code?

58
00:05:10,750 --> 00:05:18,220
Well, you can just Google Python date time format code and you'll see a list, a big list of what

59
00:05:18,220 --> 00:05:18,940
you can use.

60
00:05:19,770 --> 00:05:26,290
So let's use Year, dash and week number then.

61
00:05:26,290 --> 00:05:27,100
What's next?

62
00:05:27,370 --> 00:05:30,120
Well, next is to group the data.

63
00:05:30,790 --> 00:05:33,760
So let's say week_average.

64
00:05:35,340 --> 00:05:40,320
A new data frame equal to data.groupby

65
00:05:42,990 --> 00:05:46,800
week, so that column.

66
00:05:49,320 --> 00:05:52,260
And we want to exract the mean out of that.

67
00:05:55,130 --> 00:05:56,050
Let's see what we get.

68
00:05:58,930 --> 00:06:05,740
So it seems to be working, you see, well, Python is considering the first week zero zero, but it

69
00:06:05,740 --> 00:06:06,420
doesn't matter.

70
00:06:06,880 --> 00:06:07,720
I think it's OK.

71
00:06:08,140 --> 00:06:10,360
And so we have the average for every week.

72
00:06:12,360 --> 00:06:14,280
Again, don't forget that.

73
00:06:16,470 --> 00:06:25,350
The week column is actually the index, so that is the first week, second week and so on, and the

74
00:06:25,350 --> 00:06:26,460
rating is the column.

75
00:06:26,920 --> 00:06:28,620
So let's do the plotting now.

76
00:06:28,620 --> 00:06:33,480
Plt.plot.

77
00:06:35,890 --> 00:06:41,060
So we have to give an X and a Y, the X

78
00:06:41,120 --> 00:06:49,810
this time he's going to be week_average.index, so the weak column and

79
00:06:51,170 --> 00:06:54,470
data rating, execute.

80
00:07:01,280 --> 00:07:08,270
I got this error, X and Y must have same first dimension, 
but have shapes that and that, so it seems

81
00:07:08,270 --> 00:07:14,650
like I am using columns from different data frames.

82
00:07:14,680 --> 00:07:20,450
So this is the week average of the frame, which has 173 rows.

83
00:07:20,730 --> 00:07:24,200
This is the other data frame which has 45000 rows.

84
00:07:27,390 --> 00:07:32,310
So I meant to say week average.

85
00:07:35,450 --> 00:07:38,210
Wait, and this is the graph.

86
00:07:40,850 --> 00:07:49,310
Again, if you want to apply some sizing, you want to copy that and paste it in here, so that basically

87
00:07:49,310 --> 00:07:51,290
modifies the figure object.

88
00:07:56,010 --> 00:08:05,100
And so this is a graph, so average rating by week.
Now, you see the labels here are smashed up with

89
00:08:05,100 --> 00:08:07,800
each other and there are ways to fix that.

90
00:08:07,800 --> 00:08:16,230
But I would say it's not worth doing that here with matplotlib, because if you really want to show

91
00:08:16,230 --> 00:08:23,570
your data to people, then you want to use a more modern plotting library, such as Highcharts, which

92
00:08:23,580 --> 00:08:32,010
we are going to use in the next videos and Highcharts will show a more intelligent graph, which

93
00:08:32,010 --> 00:08:40,410
is user friendly and makes for great user experience and tries to show the information in a more efficient

94
00:08:40,410 --> 00:08:40,690
way.

95
00:08:41,100 --> 00:08:49,470
So I'll say this representation is sufficient for the purpose of using matplotlib in a Jupyter

96
00:08:49,500 --> 00:08:49,940
notebook.

97
00:08:49,960 --> 00:08:52,050
So we are just exploring data.

98
00:08:52,410 --> 00:08:59,180
We can see now we see a trend that the ratings are increasing by time.

99
00:09:00,060 --> 00:09:05,040
So maybe that was not very visible in the daily graph.

100
00:09:05,040 --> 00:09:06,300
So you see this graph here.

101
00:09:06,450 --> 00:09:12,660
This has the count currently, but we can change that to mean, execute the cell again, then execute the

102
00:09:12,660 --> 00:09:13,800
plotting sales again.

103
00:09:14,430 --> 00:09:18,460
And so we can see the daily average is here, the week average is here.

104
00:09:19,050 --> 00:09:27,840
So I believe you agree with me that it's more easy to see the trends in this graph than in this graph,

105
00:09:28,020 --> 00:09:33,140
and that is the power of downsampling the data. In the next video,

106
00:09:33,150 --> 00:09:40,200
we are going to downsample the data even further and we are going to show the average ratings by month.

107
00:09:40,770 --> 00:09:41,610
See you in the next video.

