1
00:00:00,860 --> 00:00:01,610
Hi, welcome back.

2
00:00:01,760 --> 00:00:06,620
In this video, we are going to create an average rating by month by course graph.

3
00:00:06,860 --> 00:00:12,340
As you may know, we created that graph already bought using JupYter and Matplotlib.

4
00:00:12,590 --> 00:00:13,580
So it looked like this.

5
00:00:14,900 --> 00:00:21,380
The idea is that we're going to have the month in the horizontal axis and average ratings on the vertical

6
00:00:21,380 --> 00:00:28,040
axis, and we're going to have multiple graphs in that plot area, each representing a course in our

7
00:00:28,050 --> 00:00:28,610
dataset.

8
00:00:28,850 --> 00:00:36,740
And as you may know already, a good way to create graphs using JustPy and HighCharts is to go to

9
00:00:36,740 --> 00:00:39,140
the heightcharts.com website.

10
00:00:39,140 --> 00:00:45,860
So you go under charts and series types and you find a graph that would be a good representative of our

11
00:00:45,860 --> 00:00:46,310
data.

12
00:00:46,640 --> 00:00:54,200
So I took a look in these examples and I found that I would like to use Areaspline chart, which looks

13
00:00:54,200 --> 00:00:54,650
like this.

14
00:00:54,920 --> 00:00:57,170
You see, this has multiple lines.

15
00:00:57,170 --> 00:01:02,200
So two in this case and it has this X axis with categorical data.

16
00:01:02,390 --> 00:01:03,340
So it's not numbers.

17
00:01:03,350 --> 00:01:05,060
It's Monday, Tuesday, Wednesday and so on.

18
00:01:05,450 --> 00:01:08,420
And we have this number y axis.

19
00:01:08,690 --> 00:01:16,490
So fruit units and it represents the average fruit consumption during one week for John and Jane.

20
00:01:16,700 --> 00:01:23,020
In our case, we are going to change this graph and have six or seven lines.

21
00:01:23,030 --> 00:01:25,580
I'm not sure how many courses there are.

22
00:01:26,210 --> 00:01:31,810
So one line for each course showing the average ratings along time.

23
00:01:31,880 --> 00:01:33,070
So over the months.

24
00:01:33,290 --> 00:01:36,950
So before we copy the code, let's first prepare the environment.

25
00:01:37,640 --> 00:01:41,460
So create the script and create the JustPy app and so on.

26
00:01:42,260 --> 00:01:51,200
So I'm going to name this four average rating and this would be, let's say, course month.

27
00:01:51,470 --> 00:01:58,440
So average rating by course by month dot py. 
You see that is different from average rating month.

28
00:01:58,460 --> 00:02:00,560
This is average rating by course by month.

29
00:02:01,190 --> 00:02:01,520
Right.

30
00:02:01,640 --> 00:02:07,160
Then I'm going to copy this template of JustPy, copy there.

31
00:02:09,420 --> 00:02:19,710
Hc will be the HighCharts component, which belongs to WP and options is going to be as usually charts

32
00:02:20,130 --> 00:02:31,020
definition, which has to be a variable somewhere here is equal to a string, a multiline string.

33
00:02:33,370 --> 00:02:36,040
So let us copy that JavaScript code.

34
00:02:39,100 --> 00:02:46,240
So this one in here, I'm going to copy all of that pasted in there and then just before chart, we

35
00:02:46,240 --> 00:02:49,780
leave that curly bracket before chart here.

36
00:02:49,990 --> 00:02:54,490
We leave that and delete the semicolon and the parentheses.

37
00:02:55,890 --> 00:03:00,520
And now we are ready to replace these data with our data.

38
00:03:00,720 --> 00:03:08,400
But what I suggest you do first is to try out this example to see if this is working right now without

39
00:03:08,400 --> 00:03:10,680
changing the data with your own data.

40
00:03:10,890 --> 00:03:13,980
So what I would like to do now is try to run this.

41
00:03:15,910 --> 00:03:23,970
And then visit the URL and you're going to see an error, so JsonDecodeError: Unknown identifier Highcharts.

42
00:03:24,280 --> 00:03:27,370
This happens because of this line.

43
00:03:28,360 --> 00:03:32,550
Python cannot recognize this expression here because this is JavaScript.

44
00:03:32,800 --> 00:03:38,670
So you want to delete that and only leave this string here, the color of the background color.

45
00:03:39,580 --> 00:03:49,300
And now if I stop the script and run again and reload the page, we should see the example just like

46
00:03:49,300 --> 00:03:51,680
we got it from the HighCharts website.

47
00:03:52,210 --> 00:03:58,480
Now, these two graphs are area splines, but if you want to change them to spline, you want to go

48
00:03:58,480 --> 00:04:03,370
over to your code and change that area spline to spline.

49
00:04:04,890 --> 00:04:07,170
Stop the script, run again.

50
00:04:10,720 --> 00:04:13,000
And reload and

51
00:04:14,140 --> 00:04:15,010
this is the output.

52
00:04:16,250 --> 00:04:20,660
Now we are ready to inject our own data into these data.

53
00:04:20,900 --> 00:04:26,030
So for that, we want to go over to Jupyter, where we have our analysis ready.

54
00:04:26,750 --> 00:04:29,240
Copy that code.

55
00:04:31,160 --> 00:04:39,820
Paste it there, and we also need to do the aggregations necessary for extracting the month, the courses by month,

56
00:04:40,070 --> 00:04:44,720
so which is this code here and that gives us this data frame.

57
00:04:46,520 --> 00:04:52,310
Now, let's refresh our memory, how the data frame looked like so you can either print them in here.

58
00:04:55,700 --> 00:05:01,790
And execute the script, reload the page and see what you got printed out in the terminal, or you can

59
00:05:01,790 --> 00:05:09,520
see the output from Jupyter, so whatever you prefer, I'm just going to execute the script here.

60
00:05:15,380 --> 00:05:19,760
And so we gathered data framed, printed out even before we load the...

61
00:05:20,690 --> 00:05:27,890
The Web app, because it print it out because the Web app is inside a function, so this is just in

62
00:05:27,890 --> 00:05:30,210
the global namespace, in the global scope.

63
00:05:30,500 --> 00:05:33,720
So when the script is executed, these are executed.

64
00:05:34,190 --> 00:05:41,260
If this was inside the function here, it would only be executed when we load the page.

65
00:05:41,420 --> 00:05:43,270
And yeah, this is the data frame.

66
00:05:43,850 --> 00:05:49,490
So it has this month index and it has different columns.

67
00:05:49,520 --> 00:05:51,080
So this is the first column.

68
00:05:51,380 --> 00:05:53,750
Then we have more columns which we don't see in here.

69
00:05:53,930 --> 00:05:58,050
And then the last column is this and the value of each column,

70
00:05:58,220 --> 00:06:05,060
so the name of a column, the values of that column, which are the average ratings, and we have one

71
00:06:05,060 --> 00:06:06,240
month for each row.

72
00:06:06,260 --> 00:06:09,680
So this is July of 2019.

73
00:06:10,190 --> 00:06:11,540
August 2019.

74
00:06:13,910 --> 00:06:15,770
January 2000, and so on.

75
00:06:17,540 --> 00:06:22,210
So index columns, let's see now.

76
00:06:26,140 --> 00:06:31,000
How do we inject those data into these data?

77
00:06:33,700 --> 00:06:47,770
So we see that X-axis has these categories with these weekdays, so we can do the same thing, no, down

78
00:06:47,770 --> 00:06:53,890
here we say hc.options.xAxis.

79
00:06:54,160 --> 00:06:55,750
So xAxis is the

80
00:06:57,860 --> 00:07:07,420
key of the main dictionary, so it's directly exposed in the first level, so chart titled Legend Axis,

81
00:07:07,610 --> 00:07:10,640
and then we have the second level, so categorisesniss part of axis.

82
00:07:10,890 --> 00:07:14,480
Therefore we can access it like so.

83
00:07:16,120 --> 00:07:20,390
So Catagories is a list, as you can see here.

84
00:07:20,740 --> 00:07:28,350
Therefore, we need to replace it with a list. 
List month_average_crs.

85
00:07:28,360 --> 00:07:32,890
So that is our data frame dot index.

86
00:07:33,730 --> 00:07:39,310
So that column in there, 2018 five, 2018 six and so on.

87
00:07:39,940 --> 00:07:40,270
Right.

88
00:07:40,430 --> 00:07:41,420
That was the easy part.

89
00:07:41,650 --> 00:07:43,000
Now comes the data.

90
00:07:43,930 --> 00:07:48,700
So what you need to do is you need to understand the structure of the data.

91
00:07:48,700 --> 00:07:51,520
That is the very important first step.

92
00:07:51,920 --> 00:07:54,730
And so series is

93
00:07:56,020 --> 00:08:03,500
a first level key, therefore, we can say hc.options. series, right?

94
00:08:04,030 --> 00:08:07,030
So this will give us that list.

95
00:08:08,770 --> 00:08:13,710
Then that list has... How many items does it have?

96
00:08:14,650 --> 00:08:22,540
I see a dictionary from there up to here, that's a dictionary that has this name, John pair of key

97
00:08:22,540 --> 00:08:24,880
and values, name is the key.

98
00:08:24,880 --> 00:08:25,700
John is the value.

99
00:08:25,930 --> 00:08:35,160
Then we have this other key, data and a list and the dictionary ends there, and then we have the other dictionary.

100
00:08:36,220 --> 00:08:40,050
So it's Jane now, and these are the data for Jane.

101
00:08:40,930 --> 00:08:48,530
So the list is made of dictionarys, and each dictionary represents a line in that chart.

102
00:08:49,060 --> 00:08:52,450
So that's for Jane and that's for John.

103
00:08:54,070 --> 00:08:57,280
Therefore, what we need to do is we need to have here.

104
00:08:58,310 --> 00:09:01,150
So that's got to be the name of a course.

105
00:09:01,840 --> 00:09:03,460
Let's say the first course would be

106
00:09:06,370 --> 00:09:16,420
"100 Python Exercises I" and so on, the entire title, and then data would contain a list of average ratings

107
00:09:16,420 --> 00:09:17,810
for that particular course.

108
00:09:17,830 --> 00:09:22,000
So that one, that one, that one and so on.

109
00:09:22,270 --> 00:09:29,880
And then we would have to repeat that process for each of the six or seven courses that we have here.

110
00:09:30,280 --> 00:09:31,300
How do we do that?

111
00:09:31,360 --> 00:09:33,700
Well, you could do that manually, like.

112
00:09:34,780 --> 00:09:39,220
Maybe you could inject so from serious, you would say

113
00:09:41,330 --> 00:09:49,400
the first plot, so that would give you the first dictionary, that expression is the first dictionary,

114
00:09:50,300 --> 00:09:59,930
and then we could access and name and replace that to maybe the month_average_crs.columns.

115
00:10:00,710 --> 00:10:03,140
So that would be the first column and so on.

116
00:10:03,140 --> 00:10:07,120
That be OK, but it's not an automatic process.

117
00:10:07,790 --> 00:10:12,780
So when you do manual work, that is a sign that you're not doing programming the right way.

118
00:10:13,310 --> 00:10:20,180
What we need to do instead is we need to create this entire list dynamically.

119
00:10:20,960 --> 00:10:29,550
This is a very, very crucial point now in using HighCharts, but also in knowing how to manipulate

120
00:10:29,570 --> 00:10:31,340
data structures in Python.

121
00:10:31,910 --> 00:10:37,880
So if this is an excellent practice, I'm very glad this this came out such a great practice for you.

122
00:10:38,210 --> 00:10:43,160
So let me show you how you can construct this list out of this data frame.

123
00:10:44,000 --> 00:10:50,780
So before we do that, before we assign the the list, we want to construct the list.

124
00:10:51,210 --> 00:10:55,610
So let's name this hc underscore data.

125
00:10:56,940 --> 00:10:58,760
And that is equal to.

126
00:10:58,930 --> 00:11:04,750
So we need to create a list, so that is a list.

127
00:11:04,960 --> 00:11:08,530
Therefore we write square brackets.

128
00:11:09,460 --> 00:11:12,580
The list is made of dictionaries.

129
00:11:12,850 --> 00:11:15,790
Therefore, we write curly brackets.

130
00:11:16,570 --> 00:11:21,520
And then inside each dictionary, we have a name key.

131
00:11:21,530 --> 00:11:28,380
We always have to name key and a variable for that name key.

132
00:11:28,720 --> 00:11:29,200
Why

133
00:11:29,200 --> 00:11:30,640
am I saying variable?

134
00:11:30,920 --> 00:11:33,360
Well, because name is always name.

135
00:11:33,370 --> 00:11:34,580
So it's a constant.

136
00:11:34,600 --> 00:11:37,170
That's why I'm hard coding that.

137
00:11:37,180 --> 00:11:39,520
So we should always have a name there.

138
00:11:39,670 --> 00:11:45,340
But the value of name is a variable, so it's changing every time.

139
00:11:47,270 --> 00:11:49,620
Then, what else? We have that pair.

140
00:11:49,640 --> 00:11:55,620
OK, inside the dictionary, that dictionary, so that dictionary has another pair.

141
00:11:56,030 --> 00:11:58,070
So data and the list.

142
00:11:58,190 --> 00:12:02,840
So let's write that, a comma, data,

143
00:12:03,920 --> 00:12:14,350
a string, a colon, and we have a list. That list is basically representative of each of the plots,

144
00:12:14,360 --> 00:12:17,200
each of the seven plots that we are going to have for each course.

145
00:12:18,000 --> 00:12:21,980
So which means that list is going to be these ratings.

146
00:12:22,490 --> 00:12:31,040
For example, we're going to have the first list being the first column, for example for"100 Python exercises I"

147
00:12:31,040 --> 00:12:34,940
for that course this column is a list in there.

148
00:12:34,940 --> 00:12:39,180
And then the other column after that goes the next and the next and the next.

149
00:12:39,600 --> 00:12:48,930
So one column for each course and v1 one is going to be replaced by the title of the column.

150
00:12:49,580 --> 00:12:57,800
So before I write the expression inside this list, I want to finish first the upper level list comprehension,

151
00:12:57,800 --> 00:12:59,620
because we have to list comprehension here.

152
00:12:59,780 --> 00:13:02,930
We have the first level one, so let's call it the upper level.

153
00:13:03,830 --> 00:13:06,980
Here we are constructing dictionaries.

154
00:13:07,580 --> 00:13:09,440
So it's a list made of dictionaries.

155
00:13:10,140 --> 00:13:10,850
So

156
00:13:10,850 --> 00:13:22,100
that means we should say here for v1 in month_average_crs.columns.

157
00:13:23,420 --> 00:13:32,860
So what that is, is a list containing seven strings, one string for each course title.

158
00:13:32,870 --> 00:13:36,740
So it's going to be a list containing "100 Python exercises I:

159
00:13:36,740 --> 00:13:41,960
Evaluate and improve your skills", then the other title of the other course and so on.

160
00:13:41,960 --> 00:13:44,180
It's a list of these titles.

161
00:13:44,570 --> 00:13:53,150
So what we're doing is we are injecting here a title, the first title of that list in the first iteration,

162
00:13:53,330 --> 00:13:58,610
then the next title of that month_average_crs.columns lists in the next iteration.

163
00:13:58,610 --> 00:14:05,000
And so one, we're going to have like seven iterations or six or whatever is the number of courses that

164
00:14:05,000 --> 00:14:05,540
we have.

165
00:14:05,840 --> 00:14:06,130
Right.

166
00:14:06,140 --> 00:14:13,430
But then, um, inside here we have a nested iteration, a nested list comprehension.

167
00:14:13,430 --> 00:14:15,550
Sorry, which is this one here.

168
00:14:16,070 --> 00:14:28,850
And that will be again, let's write another variable such as v2 for v2 in month average 

169
00:14:28,850 --> 00:14:29,740
crs.

170
00:14:30,090 --> 00:14:30,570
Hmm.

171
00:14:30,650 --> 00:14:33,860
So what is this going to be?

172
00:14:33,920 --> 00:14:36,830
Well, it has to be this column.

173
00:14:36,830 --> 00:14:37,240
Right.

174
00:14:38,420 --> 00:14:40,540
So which which one is this column?

175
00:14:40,580 --> 00:14:41,990
What's the name of this column?

176
00:14:42,200 --> 00:14:45,870
Well, for this course is that is the name of the column.

177
00:14:46,280 --> 00:14:50,540
So here we need to write down the name of the column.

178
00:14:50,690 --> 00:14:53,870
You see, this is the Pandas syntax.

179
00:14:55,120 --> 00:14:55,880
So you write

180
00:14:55,900 --> 00:15:04,930
the name of the column, but we have different column names, so what we enter here is v1.

181
00:15:06,160 --> 00:15:13,880
Because one is going to be the current title in the upper level list comprehension.

182
00:15:14,050 --> 00:15:22,360
So basically what Python will do is while it is iterating the first title, it's going to stop in that

183
00:15:22,360 --> 00:15:29,560
upper level, first iteration, and it's going to iterate through that particular column, through these

184
00:15:29,590 --> 00:15:33,010
ratings for "100 Python exercises I".

185
00:15:33,550 --> 00:15:36,220
And it will construct that list. Once it finishes

186
00:15:36,220 --> 00:15:41,680
with all these numbers, it constructs that list,

187
00:15:41,900 --> 00:15:49,070
then it goes to the other to the second iteration in the first level iteration.

188
00:15:49,300 --> 00:15:55,620
So in the second iteration of the other columns, which is going to be the next course and so on.

189
00:15:55,900 --> 00:15:58,090
I hope it's clear.

190
00:15:58,840 --> 00:15:59,270
Right.

191
00:15:59,350 --> 00:16:08,470
So what's left to do is we need to assign this new list so to inject it to our data because we created

192
00:16:08,470 --> 00:16:15,390
it, but we haven't connected it to hc, so it's just a variable floating there in our script.

193
00:16:15,790 --> 00:16:17,590
No connections to the other parts.

194
00:16:18,110 --> 00:16:26,610
So what we need to do is we said that that one that list is that list.

195
00:16:27,280 --> 00:16:36,190
So we need to assign it to series, right? Hc.options.series equal to hc.data.

196
00:16:37,000 --> 00:16:41,050
Right. Now I'm going to stop the script and run.

197
00:16:42,630 --> 00:16:46,140
And visit the app, and this is the output.

198
00:16:47,960 --> 00:16:57,440
Here is X-axis with the month, and this is our ratings from two point five up to five point five,

199
00:16:58,460 --> 00:17:00,410
the maximum is somewhere in here.

200
00:17:01,160 --> 00:17:06,740
And so you can hover your mouse there and you see the ratings for each course.

201
00:17:07,400 --> 00:17:07,850
Right.

202
00:17:08,270 --> 00:17:11,170
The legend here is floating.

203
00:17:11,570 --> 00:17:14,690
So if you want to change that, you can go back to your code

204
00:17:15,760 --> 00:17:23,050
and locate the legend here are so you have some parameters here, which, again, experiment with but

205
00:17:23,620 --> 00:17:26,080
if I set floating to false,

206
00:17:30,980 --> 00:17:35,220
that should change the view of the graph.

207
00:17:35,240 --> 00:17:37,970
So I think this is more clear now.

208
00:17:38,840 --> 00:17:47,090
I wouldn't agree that this type of graph is the best representation for these kind of data because we

209
00:17:47,090 --> 00:17:51,070
have multiple lines here, but they are so close to each other.

210
00:17:51,500 --> 00:17:56,210
So there's not much variations between one course and the others.

211
00:17:57,500 --> 00:18:04,550
I think these splines would be more representative if we had instead of averages,

212
00:18:04,550 --> 00:18:11,240
we have the counts of the reviews of the ratings left for each course, which means we have to

213
00:18:11,240 --> 00:18:16,040
go to the processing section and say here count instead of mean.

214
00:18:16,700 --> 00:18:17,720
And let me try that.

215
00:18:22,850 --> 00:18:23,170
Hmm.

216
00:18:23,520 --> 00:18:30,590
So, yeah, it looks better I think. We have more variation, more difference between this plots and

217
00:18:30,590 --> 00:18:38,240
the other ones, although the other ones are still not as visible since this plot here is taking

218
00:18:38,240 --> 00:18:42,820
a lot of the range of the y axis range.

219
00:18:43,210 --> 00:18:43,700
Hmm.

220
00:18:43,760 --> 00:18:49,940
Therefore, what I'm going to do next, in the next vide is I'm going to use another type of graph to

221
00:18:49,940 --> 00:18:57,680
represent the same exact same data, average ratings because by months, but using a stream graph.

222
00:18:57,900 --> 00:18:59,170
I'll talk to you in the next video.

