1
00:00:00,180 --> 00:00:01,080
Hi, welcome back.

2
00:00:01,650 --> 00:00:07,920
When you're doing data analysis, the first thing you want to do is you want to familiarize yourself with

3
00:00:07,920 --> 00:00:14,550
your data using the tool, the data analysis tool that you have chosen to use for your particular

4
00:00:14,550 --> 00:00:15,120
project.

5
00:00:15,300 --> 00:00:17,680
In this case, the tool is Python.

6
00:00:18,270 --> 00:00:24,710
So what we're going to do in this video is we are going to load the data into Python using a Jupyter

7
00:00:24,720 --> 00:00:32,630
notebook and then will look at those data and extract some very basic information about the data, such

8
00:00:32,640 --> 00:00:39,810
as looking at what column names we have, how many rows we have and other simple attributes of our

9
00:00:39,810 --> 00:00:40,200
data.

10
00:00:41,040 --> 00:00:41,970
So let's do that.

11
00:00:42,340 --> 00:00:49,590
I'd like to let you know that I created an empty folder named review_analysis, and I

12
00:00:49,830 --> 00:00:57,650
put the reviews.csv file in that folder.
You can find this reviews.csv file

13
00:00:57,660 --> 00:00:59,210
attached to this lecture.

14
00:00:59,310 --> 00:01:00,870
So as e lecture resource.

15
00:01:01,170 --> 00:01:04,980
So please download that and place it in a folder just like I did.

16
00:01:05,200 --> 00:01:14,130
And then you should be able to locate that folder from the Jupyter homepage, which is localhost 

17
00:01:14,210 --> 00:01:16,050
8888/tree.

18
00:01:16,680 --> 00:01:20,740
So I created that folder directly in my user's folder.

19
00:01:20,940 --> 00:01:22,320
So this is the folder.

20
00:01:22,320 --> 00:01:25,740
I can click that and here is the reviews.csv file.

21
00:01:26,730 --> 00:01:32,520
In your case, you can create this wherever you want and then locate it using this directory tree here.

22
00:01:34,980 --> 00:01:44,130
While I am in this folder from Jupyter, I can go to the new dropdown list and go to Python three to

23
00:01:44,130 --> 00:01:46,250
create a new Jupyter notebook.

24
00:01:47,220 --> 00:01:54,870
So the Jupyter notebook has been created so I can rename this to something else, the name of the Jupyter

25
00:01:54,870 --> 00:01:55,320
notebook.

26
00:01:55,890 --> 00:01:57,900
Let's keep it simple and see reviews.

27
00:01:59,200 --> 00:02:08,260
Rename, and that will get a .pynb extension, let me make some more room here so that you

28
00:02:08,260 --> 00:02:12,990
see more code, I'm going to toggle the header of.

29
00:02:13,000 --> 00:02:14,050
Now, let's start coding.

30
00:02:15,400 --> 00:02:22,360
The very first thing we want to do, of course, is import pandas, the library that is used to perform

31
00:02:22,360 --> 00:02:24,070
data analysis with Python.

32
00:02:24,580 --> 00:02:31,720
And then we want to create a variable which is going to hold the data that is equal to pandas.read.

33
00:02:31,870 --> 00:02:40,390
Since we are working with a Csv file the method we want to use out of the pandas library is read_csv.

34
00:02:40,390 --> 00:02:42,880
In parentheses

35
00:02:42,880 --> 00:02:51,050
we want to pass in single quotes or double quotes, whatever you like the path to the csv file.

36
00:02:51,090 --> 00:02:56,530
Now, if you don't type in the name correctly, you're going to get an error.

37
00:02:57,970 --> 00:03:05,590
So let me try to execute this sale using control enter if you are on windows or command enter if you

38
00:03:05,590 --> 00:03:10,690
are on Mac and as I warned you, I got an error.

39
00:03:10,720 --> 00:03:15,280
It says no such file or directory because I mistyped the file.

40
00:03:15,400 --> 00:03:18,610
So 'reviews'. This time

41
00:03:18,610 --> 00:03:20,700
if I execute, I don't get an error.

42
00:03:20,710 --> 00:03:28,510
That means data was loaded successfully and I can press 'escape b enter'

43
00:03:30,200 --> 00:03:35,660
and call the data variable and control, enter again to see that data frame.

44
00:03:39,410 --> 00:03:46,010
So we can see that we have one, two, three, four columns, and we also have this index column here

45
00:03:46,010 --> 00:03:53,160
added by Pandas automatically and basically is a range of numbers starting from zero.

46
00:03:53,180 --> 00:04:00,140
So that is the first row of our data, this one in here that has this index of zero.

47
00:04:01,760 --> 00:04:03,500
And it ends at the last row.

48
00:04:04,490 --> 00:04:12,350
You can see that this is the first row, second, third, fourth, fifth row and then Jupyter is not displaying

49
00:04:12,350 --> 00:04:21,550
the rows after row five because there are a lot of rows and it would be impractical to see them here.

50
00:04:21,680 --> 00:04:30,920
However, you get to see the last five rows over the data frame, and that gives you an overview

51
00:04:30,920 --> 00:04:31,900
of the data frame.

52
00:04:32,510 --> 00:04:42,310
However, what I like to do instead is just print out the heads of the frame, which is the first rows

53
00:04:42,320 --> 00:04:42,880
only.

54
00:04:43,790 --> 00:04:45,620
So the first five rows.

55
00:04:46,130 --> 00:04:48,320
That gives you a more compact view.

56
00:04:48,500 --> 00:04:53,420
It gives you an idea what columns you have and what kind of rows you have also.

57
00:04:53,930 --> 00:04:56,900
So it's good to have the head of the data frame displayed in here.

58
00:04:57,170 --> 00:05:04,220
Then we can press 'escape b' and create a new cell here, enter to write some other code.

59
00:05:05,570 --> 00:05:11,410
You can get the shape of the data frame by accessing the shape property.

60
00:05:11,630 --> 00:05:13,070
So that is a method.

61
00:05:14,360 --> 00:05:22,330
With parentheses, that is a propery. It doesn't need parentheses, and then you get to the shape of the data

62
00:05:22,550 --> 00:05:27,650
frame, which is basically the number of rows and the number of columns.

63
00:05:29,940 --> 00:05:32,910
You see one, two, three, four columns.

64
00:05:34,900 --> 00:05:46,840
You might also want to... B enter. Display the columns that your data fraim has, even though we

65
00:05:46,840 --> 00:05:47,680
have them there.

66
00:05:48,580 --> 00:05:55,960
This is yet another way to see the names of the columns by accessing the columns' property. Then, next

67
00:05:55,960 --> 00:05:57,880
usually when you are working with data,

68
00:05:58,660 --> 00:06:05,220
you have some specific columns that you are interested about, which could be one column or more.

69
00:06:05,230 --> 00:06:12,970
In this case, we might be interested to see an overview of the rating column, to see what the minimum

70
00:06:13,150 --> 00:06:21,610
rating is and what the maximum rating is and the distribution of those ratings and have them as a graph

71
00:06:21,610 --> 00:06:26,450
here, displayed here so that we can have a better understanding of our data.

72
00:06:27,190 --> 00:06:32,030
So I'm going to do escape B Enter and data.hist.

73
00:06:32,470 --> 00:06:35,700
So this would be a histogram. With parentheses.

74
00:06:37,300 --> 00:06:41,540
We are interested to see the distribution of the ratings.

75
00:06:42,430 --> 00:06:46,030
Therefore I enter the rating column here as a string.

76
00:06:48,310 --> 00:06:51,550
Execute and we get this graph.

77
00:06:52,660 --> 00:06:58,300
Let me explain to you what the graph means. What this means is that, for example, let's start from the right.

78
00:06:59,560 --> 00:07:01,930
This bar here, this first bar

79
00:07:03,110 --> 00:07:08,780
means that we have around 24 000

80
00:07:09,760 --> 00:07:18,010
five star reviews, ratings, you see, for example, we have this this is a 5.0 star rating,

81
00:07:18,220 --> 00:07:24,300
so we have around 24 000 five star ratings in the whole data frame.

82
00:07:24,340 --> 00:07:29,750
And in total, we have 45 000 rows in total.

83
00:07:30,160 --> 00:07:33,130
Then the next, this bar here.

84
00:07:35,320 --> 00:07:41,950
These are 4.5 ratings like that one in there.

85
00:07:43,830 --> 00:07:51,420
And we have around, let's say, 7000 of those, then we have four star ratings

86
00:07:52,420 --> 00:07:59,470
around 9000, perhaps. 3.5 star ratings in here.

87
00:08:00,960 --> 00:08:11,700
This is 3 star ratings, it's about 2000, then we have 2.5 star ratings this one in here, 2 star

88
00:08:11,700 --> 00:08:12,450
ratings.

89
00:08:13,800 --> 00:08:22,920
1.5 star ratings, that's the lowest number among all the ratings, so people don't leave

90
00:08:22,920 --> 00:08:25,200
a lot of 1.5 star ratings.

91
00:08:25,500 --> 00:08:28,950
And we also have this one star ratings in here.

92
00:08:29,220 --> 00:08:30,910
It's not the best graph.

93
00:08:31,200 --> 00:08:36,860
I don't like it personally, but it's a quick way so that you know how your data are distributed.

94
00:08:37,260 --> 00:08:38,450
You know, that's OK.

95
00:08:38,460 --> 00:08:45,720
We have data from one to five and five is most occurring value in your data.

96
00:08:46,910 --> 00:08:53,280
And that's about this lecture list, so you can get an overview of your data frame. In the next lecture,

97
00:08:53,300 --> 00:09:00,980
we are going to zoom in into our data frame to actually be able to select particular rows of particular

98
00:09:00,980 --> 00:09:05,400
slices of our data frame to see individual values.

99
00:09:05,750 --> 00:09:13,820
So in other words, we are going to use Python to navigate through our data and select particular sections

100
00:09:13,820 --> 00:09:15,930
of the data and display them.

101
00:09:16,250 --> 00:09:17,450
So I'll see you in the next video.

