1
00:00:00,960 --> 00:00:07,470
Hello and welcome to this new video, which is going to be a bit different from the other ones.

2
00:00:07,470 --> 00:00:14,610
In the previous videos, I showed you some techniques to filter out data from a data frame, and that was

3
00:00:14,610 --> 00:00:22,440
basically accessing columns and rows and cells and also applying conditions when accessing these columns

4
00:00:22,440 --> 00:00:25,170
or rows or cells of a data frame.

5
00:00:25,650 --> 00:00:32,010
However, what we have done so far is just extracting data out of data.

6
00:00:32,760 --> 00:00:34,790
This is still not data analysis

7
00:00:34,830 --> 00:00:41,250
because the goal of data analysis is to turn data are into information.

8
00:00:42,560 --> 00:00:52,220
That is what data analysis is about, but this is still not information. For example, let's say we

9
00:00:52,220 --> 00:00:58,430
want the ratings above four and we got this data frame, but it's still a data frame.

10
00:00:58,430 --> 00:01:00,350
It's not human friendly.

11
00:01:00,380 --> 00:01:02,030
It's not information.

12
00:01:02,390 --> 00:01:08,660
Information would be if we instead got the average of the ratings above four.

13
00:01:08,810 --> 00:01:11,410
So we would end up with one single number.

14
00:01:11,630 --> 00:01:14,510
And that is information because it's human readable.

15
00:01:14,510 --> 00:01:15,860
It tells us something.

16
00:01:16,130 --> 00:01:23,080
Or maybe we could create a plot, a graph where we could see the ratings over time.

17
00:01:23,660 --> 00:01:24,740
So we see the graph.

18
00:01:24,740 --> 00:01:30,390
We can and we can see a trend and we can learn some information from that graph.

19
00:01:30,650 --> 00:01:34,130
So turning the other into information is what we are going to do next.

20
00:01:34,520 --> 00:01:39,680
And for this particular video, we are going to answer a series of questions.

21
00:01:40,160 --> 00:01:45,110
In other words, we are going to extract some information such as we are going to give the average rating,

22
00:01:45,350 --> 00:01:47,410
the average rating for a particular course.

23
00:01:47,420 --> 00:01:48,140
And so on.

24
00:01:49,410 --> 00:01:56,610
Now, you should be able to actually write the code before me, because you already know, for example,

25
00:01:56,610 --> 00:02:03,120
how to extract the rating column, and I also gave you some clues that you could use a mean method

26
00:02:03,120 --> 00:02:06,120
to get the average rating of that column.

27
00:02:06,600 --> 00:02:08,350
So let's start with the first information.

28
00:02:08,370 --> 00:02:16,710
The average rating for all the courses all times.
That would be data rating dot mean.

29
00:02:18,900 --> 00:02:20,160
And that is the average rating.

30
00:02:20,400 --> 00:02:22,170
So just like we had here.

31
00:02:24,390 --> 00:02:26,340
In selecting a column

32
00:02:28,960 --> 00:02:35,710
data rating, that's what we do here as well, so that gives us the column and that gives us the rating.

33
00:02:35,950 --> 00:02:38,680
Right, next average rating for a particular course.

34
00:02:40,120 --> 00:02:49,230
This time we have to apply a condition, which is data course name is equal...

35
00:02:50,170 --> 00:02:56,920
So double assignment operator is equal to one of the courses.

36
00:02:58,170 --> 00:03:07,830
Well, let's say "The Python Mega Course: Build 10 Real World Applications".

37
00:03:08,910 --> 00:03:14,880
Now, if you miss a single letter in the string, you're not going to get the expected result.

38
00:03:15,120 --> 00:03:19,080
Let's say you wrote a small a instead of uppercase A.

39
00:03:20,220 --> 00:03:22,400
And you get an empty data frame.

40
00:03:22,740 --> 00:03:25,530
If I change the lowercase a to uppercase,

41
00:03:25,680 --> 00:03:27,810
then we get the filtered data frame.

42
00:03:28,320 --> 00:03:32,640
And out of that, we want the rating column.

43
00:03:33,810 --> 00:03:35,930
So those are the ratings for "The Python

44
00:03:35,940 --> 00:03:45,510
Mega Course" and mean is the average of "The Python Mega Course", the average rating, which is a bit higher

45
00:03:45,510 --> 00:03:47,820
than the average rating of all the courses.

46
00:03:49,310 --> 00:03:49,660
Right.

47
00:03:49,700 --> 00:03:52,060
Average rating for a particular period.

48
00:03:53,700 --> 00:03:57,900
Data, I'm just going to copy this

49
00:03:59,600 --> 00:04:09,170
and paste it here, so we have this first condition, and that second condition, so we're talking about

50
00:04:09,170 --> 00:04:16,430
this period, you can change that for another period, let's say the entire 2020 from 1st of January

51
00:04:16,670 --> 00:04:19,460
to the 31st of December.

52
00:04:20,750 --> 00:04:29,480
Now, you can split this expression and still get the same output by entering, by pressing enter after

53
00:04:29,480 --> 00:04:35,390
the end operator and you still get the filtered data frame, out of that

54
00:04:35,390 --> 00:04:36,680
ee get to rating.

55
00:04:39,380 --> 00:04:48,800
And out of that, we get the mean, and that is the mean for 2020. Next average rating for a particular

56
00:04:48,800 --> 00:04:56,420
period, for a particular cause, I'm going to cope with this by pressing CC and click here

57
00:04:56,420 --> 00:05:01,780
and press V and then after this condition.

58
00:05:01,790 --> 00:05:08,960
So one parenthesis to parentheses here, I'm going to add the and operator.

59
00:05:10,200 --> 00:05:18,120
And press enter, press enter again, and here I'm going to add the other condition, which is this

60
00:05:18,120 --> 00:05:24,910
one in here, the course is equal to that particular string, "The Python Mega Course".

61
00:05:26,970 --> 00:05:27,630
Execute.

62
00:05:29,820 --> 00:05:36,420
We do have a mismatch of parentheses, you see that this square bracket is highlighted in red, that

63
00:05:36,420 --> 00:05:45,810
means we should remove it, execute again, and this is the rating for "The Python Mega Course" for 2020.

64
00:05:46,330 --> 00:05:49,590
So this is how you get information out of your data.

65
00:05:51,210 --> 00:05:55,710
Let's carry on with average of uncommented ratings.

66
00:05:58,180 --> 00:06:02,050
So uncommented ratings, you know that.

67
00:06:05,420 --> 00:06:12,260
Let me put that head of the frame here just for convenience, so you know that we have a comment column

68
00:06:12,260 --> 00:06:18,080
here, which could be an NaN, which means not a number.

69
00:06:18,470 --> 00:06:26,330
So it's no value or it could be a string, something that the student wrote as a review for a particular

70
00:06:26,330 --> 00:06:26,810
course.

71
00:06:28,860 --> 00:06:35,820
We want to get the average of the ratings that have no comments, so to do that, you say data,

72
00:06:36,120 --> 00:06:48,510
the condition data comment is no, this is a method, actually, and it gives us a filtered data frame

73
00:06:48,510 --> 00:06:53,100
with only the rows that have the comment value

74
00:06:53,490 --> 00:06:57,780
as.NaN and the opposite of that is not null.

75
00:06:58,500 --> 00:07:01,470
Then you get the rules with comment.

76
00:07:02,550 --> 00:07:04,560
But we want is null this time.

77
00:07:05,010 --> 00:07:10,560
And out of that we want that rating column and the mean.

78
00:07:12,570 --> 00:07:16,410
That is the mean, the opposite of that would be, of course

79
00:07:20,160 --> 00:07:28,920
not null, and we see that the average rating for ratings with comments is higher than the average

80
00:07:28,920 --> 00:07:35,730
of ratings without a comment, which I think is normal, people who like the course maybe tend to also

81
00:07:35,730 --> 00:07:39,110
write something to express their gratitude or I don't know.

82
00:07:39,300 --> 00:07:42,390
But this is telling us something, right?

83
00:07:42,690 --> 00:07:45,180
Number of uncommented ratings.

84
00:07:46,050 --> 00:07:47,570
Well, it's easy.

85
00:07:47,580 --> 00:07:49,860
Just copy that, paste it here,

86
00:07:49,860 --> 00:07:54,270
and instead of mean, you say count and this is the count.

87
00:07:55,050 --> 00:07:56,310
The opposite of that would be.

88
00:07:58,810 --> 00:08:11,770
Not null, so ratings with comment count, and that's plus that, I'm sure gives 45000, which is

89
00:08:11,770 --> 00:08:15,230
a total number of rows of ratings.

90
00:08:15,250 --> 00:08:18,960
So ratings with no comment, ratings with comment.

91
00:08:19,720 --> 00:08:26,440
Next number of comments containing a certain word.
In the data frame in the comments,

92
00:08:27,320 --> 00:08:32,230
let's say some students talk about the accents of the instructor.

93
00:08:32,410 --> 00:08:39,790
What does the average of the ratings that contain accent?
Are people complaining about the accent and

94
00:08:39,790 --> 00:08:43,040
how many people are complaining about the accent of the instructor?

95
00:08:43,340 --> 00:08:45,010
Let's get that information.

96
00:08:47,340 --> 00:08:59,760
Again, the condition is that data comment, the string of that comment contains the word accent.

97
00:09:02,230 --> 00:09:04,190
What we get is an error.

98
00:09:05,280 --> 00:09:14,580
The problem is that Python cannot mask with non-boolean array containing null values, so Python can

99
00:09:14,580 --> 00:09:23,590
not search in those null values in comments, cannot search for a string in these types of data.

100
00:09:24,330 --> 00:09:33,780
So we want to tell Python through the na argument set to false. In that case Python

101
00:09:33,780 --> 00:09:42,780
Will ignore those new values and it will only search for accents in the comments with a string, with

102
00:09:42,780 --> 00:09:43,480
a value.

103
00:09:44,310 --> 00:09:46,050
And this is the filtered data frame.

104
00:09:46,140 --> 00:09:49,920
So all these comments are actually mentioning the word accent.

105
00:09:52,870 --> 00:09:54,040
How many of them?

106
00:09:59,000 --> 00:10:09,050
Those are the ratings only, the count of the ratings is 77, so 77 out of 45000 total.

107
00:10:11,420 --> 00:10:21,320
Of course, the average would be simply that, but changing that count to mean and it's a low rating,

108
00:10:21,320 --> 00:10:28,400
as I was expecting, so because the courses are taken from students, from different countries, some

109
00:10:28,400 --> 00:10:32,240
of them may find the accent of the instructor unpleasant.

110
00:10:32,540 --> 00:10:38,440
So maybe they are leaving a negative comment and a negative rating along the comment.

111
00:10:38,720 --> 00:10:40,280
And that concludes

112
00:10:41,230 --> 00:10:43,240
our lecture here.

113
00:10:45,460 --> 00:10:50,980
Thanks for following! In the next videos, we are going to do some plotting, which is a lot of fun

114
00:10:50,980 --> 00:10:57,700
in my opinion, and it's a beautiful way to present information to the public, to the user.

115
00:10:58,060 --> 00:10:58,390
Thanks.

116
00:10:58,390 --> 00:10:59,340
And I'll talk to you later.

