1
00:00:00,510 --> 00:00:01,300
All right.

2
00:00:01,320 --> 00:00:04,910
Now that's you know all to load data in Python via pandas.

3
00:00:05,220 --> 00:00:10,710
And you know how to do that through using different data sources as you saw, csv, Json , text files and

4
00:00:10,710 --> 00:00:19,680
excel and now you want to understand and learn how to manipulate these data frames and by manipulation

5
00:00:19,710 --> 00:00:27,420
what I mean is deleting rows and columns from your data frame and adding new rows and columns and also

6
00:00:27,420 --> 00:00:30,330
modifying existing rows and columns.

7
00:00:30,390 --> 00:00:34,350
So that's what you're going to learn throughout these lectures.

8
00:00:34,440 --> 00:00:41,730
But first of all I'd like you to understand how data frames are indexed and with indexing I mean you

9
00:00:41,730 --> 00:00:45,930
know we have this data frame here and this can be a big one.

10
00:00:45,970 --> 00:00:51,320
So this happens to be a shorter one with only six rows but if you have big data frames with lots

11
00:00:51,360 --> 00:00:58,530
of columns and rows, then you may want to extract information out of the data frame and to extract information

12
00:00:58,530 --> 00:01:02,840
you need to have like a coordinate system.

13
00:01:02,940 --> 00:01:06,090
In that data frame like an embedded coordinate system.

14
00:01:06,210 --> 00:01:09,150
So that if you want to access let's say,

15
00:01:09,150 --> 00:01:14,880
so these two rows here, this portion here you want to know how to do that.

16
00:01:15,090 --> 00:01:16,710
So that's what you're going to learn now.

17
00:01:16,830 --> 00:01:19,890
How data frames are indexed and how you can slice them.

18
00:01:19,890 --> 00:01:23,920
So let's try to extract that portion of the data frame.

19
00:01:23,940 --> 00:01:30,630
There might be different ways to access that portion of the data frame. The first way is to use label

20
00:01:30,630 --> 00:01:32,130
based indexing.

21
00:01:32,130 --> 00:01:35,040
The other way is to use position based indexing.

22
00:01:35,040 --> 00:01:40,060
So your data frame has column labels and index labels.

23
00:01:40,290 --> 00:01:46,650
So now you can use labels from your index column and labels from your header, your column names

24
00:01:47,040 --> 00:01:54,770
to access portions of your data frame. With label indexing you want to use the lock in there, so the lock

25
00:01:54,780 --> 00:02:00,810
method and then you pass square brackets in there and then that gets two elements.

26
00:02:00,870 --> 00:02:05,610
And the first element could be a range of the index column.

27
00:02:06,030 --> 00:02:08,790
So we're talking about labels about strings.

28
00:02:08,790 --> 00:02:20,060
So you'll have to pass "735 Dolores St", and then a range so with a colon there "332

29
00:02:20,160 --> 00:02:30,180
Hill St", and then from country to ID, execute that.

30
00:02:30,320 --> 00:02:31,770
And yeah, this is our portion.

31
00:02:31,980 --> 00:02:36,290
So when you use labels you're including the first label that you pass there.

32
00:02:36,540 --> 00:02:38,340
And the last one as well.

33
00:02:38,610 --> 00:02:44,090
So everything between those and the like here country and employees is included as well.

34
00:02:44,100 --> 00:02:55,170
But ID also, and of course similarly, almost similarly you can access you know, single cells from your data

35
00:02:55,170 --> 00:02:57,320
frame just like that.

36
00:02:57,360 --> 00:03:06,180
So the intersection between this index label and this column name is USA which should be this one

37
00:03:06,180 --> 00:03:06,380
here.

38
00:03:06,870 --> 00:03:16,470
If you want the USA-s, then you just pass everything there and you get everything here which of course

39
00:03:16,500 --> 00:03:21,240
if you want you can convert it to a list.

40
00:03:21,720 --> 00:03:30,530
So a simple list using the python built in function which is a list and that's about a label based indexing.

41
00:03:30,530 --> 00:03:37,890
Now this is not the common way to access, to extract data from the data frame. More common could be to

42
00:03:37,920 --> 00:03:43,210
access the database on indexing not based on labels.

43
00:03:43,630 --> 00:03:56,060
So to do that you do df7 and instead of lock you do ilock, and that again expects two items.

44
00:03:56,100 --> 00:04:05,520
So the first would be the range of your indexes. Actually let me print all the data frame here so that

45
00:04:05,550 --> 00:04:10,200
you can refer to that.

46
00:04:10,530 --> 00:04:16,960
So let me excess from Dolores to 23rd street.

47
00:04:17,020 --> 00:04:21,490
That would be one, two, three.

48
00:04:21,510 --> 00:04:22,330
I believe, yeah.

49
00:04:23,250 --> 00:04:30,880
And also from country to ID, so again one, two, three.

50
00:04:34,290 --> 00:04:36,630
And yeah, you can see the difference now.

51
00:04:37,010 --> 00:04:43,590
You know the I.D. wasn't included there and either was 23rd street because this is as you

52
00:04:43,590 --> 00:04:46,410
do with lists, this is upper bound exclusive.

53
00:04:46,410 --> 00:04:54,270
So with python lists three is not included in the slice, but with the labels that, the last item of the

54
00:04:54,270 --> 00:04:57,500
range was included in the slice.

55
00:04:57,510 --> 00:05:03,570
So in this case you want to pass 4 there, and 4 there and that's how you get your portion.

56
00:05:03,570 --> 00:05:07,050
And of course similarly you can do things like that.

57
00:05:07,800 --> 00:05:11,160
So you get all the rows or only one of them

58
00:05:14,980 --> 00:05:20,430
so that be a row with index three which is this one but only four columns

59
00:05:20,440 --> 00:05:25,480
country, employees, and ID. So USA, 10, and 4.

60
00:05:25,490 --> 00:05:26,490
Alright.

61
00:05:26,620 --> 00:05:28,150
That is position based index.

62
00:05:28,500 --> 00:05:33,460
And yeah that's what I wanted to teach you about data frame indexing and slicing.

63
00:05:33,460 --> 00:05:35,000
And I'll talk to you in the next lecture.

