1
00:00:00,199 --> 00:00:06,509
So as you see it's quite a simple one
and intentionally I tried to find a

2
00:00:06,509 --> 00:00:12,530
simple web page for you so here we go.
I didn't want to distract you with lots of

3
00:00:12,530 --> 00:00:19,380
content for now. Later you'll be able
to grab information from a big website

4
00:00:19,380 --> 00:00:29,070
with lots of data so for now let's try
to grab, let's say we want to extract the

5
00:00:29,070 --> 00:00:36,210
names of the cities from this page, so if
you want to follow me please typing this

6
00:00:36,210 --> 00:00:44,520
address on your address bar so with the
dot HTML at the end, and so we've got

7
00:00:44,520 --> 00:00:49,739
only three cities here that we will be
extracting by the code that we will

8
00:00:49,739 --> 00:00:56,730
write will work with any number of rows here.
I'll be using the iPython notebook

9
00:00:56,730 --> 00:01:02,489
or the Jupiter notebook as it is called
now so it was renamed to Jupiter

10
00:01:02,489 --> 00:01:08,930
notebook, so right right click and
open your command line.

11
00:01:08,930 --> 00:01:21,180
Jupiter notebook, and I'll create a Python 3,
a notebook, great, so the first thing you

12
00:01:21,180 --> 00:01:30,259
want to do is you want to load this
source code in Python, and the way to do that

13
00:01:30,259 --> 00:01:37,740
is by using the requests library. So if
you don't have that installed you can

14
00:01:37,740 --> 00:01:47,189
just go ahead and install it with pip
install requests. Just like that. I have it

15
00:01:47,189 --> 00:01:54,899
already so already satisfied, but the
process is very easy so you already know

16
00:01:54,899 --> 00:02:00,149
how to install packages with pip.
You'll also need the BeautifulSoup

17
00:02:00,149 --> 00:02:06,689
library, so to install that you need to
say pip install again and not BeautifulSoup,

18
00:02:06,689 --> 00:02:15,360
but bs4 so which stands for BeautifulSoup 4,
so that's the latest version of

19
00:02:15,360 --> 00:02:22,860
BeautifulSoup, and then so you want
to import quest and so the first thing

20
00:02:22,860 --> 00:02:28,290
you want to do is load the source code.
And then we start looking for HTML tags

21
00:02:28,290 --> 00:02:33,989
and extracting elements from those tags.
Now let me import BeautifulSoup as well.

22
00:02:33,989 --> 00:02:42,540
So from bs4 import
BeautifulSoup,

23
00:02:42,540 --> 00:02:48,510
so that's the syntax you're importing
the BeautifulSoup cass from bs4.

24
00:02:48,510 --> 00:02:53,400
If you are on Python 2 this should be
slightly different, so you want to import

25
00:02:53,400 --> 00:02:58,370
BeautifulSoup directly like this.
Okay, alt enter and go to the next line.

26
00:03:02,480 --> 00:03:10,069
So to load a webpage it's good to create a
variable so you can load the web page

27
00:03:10,069 --> 00:03:18,900
source code to this variable so R equals
to requests.get, so the get method.

28
00:03:18,900 --> 00:03:23,549
So you point to the library and then to
the get method and all you need to pass

29
00:03:23,549 --> 00:03:29,639
here is the URL of the web page that you
want to load, so in this case

30
00:03:29,639 --> 00:03:37,440
pythonhow.com.example.html, so don't
forget the HTML. This is just a static

31
00:03:37,440 --> 00:03:48,859
web page so you should pass HTML there.
Now this should create a requests object

32
00:03:48,859 --> 00:03:56,160
so we're still not there and what you
want to do is grab the content from this

33
00:03:56,160 --> 00:04:02,639
requests data type and maybe store it in
another variable, so the content store it in

34
00:04:02,639 --> 00:04:12,629
in the C variable like that and if you
want to check now what this is. You'll

35
00:04:12,629 --> 00:04:20,030
see that this is a bytes data type,
and you can print it if you want

36
00:04:20,380 --> 00:04:27,010
even though this one doesn't look very
nice this is actually the source code

37
00:04:27,010 --> 00:04:36,980
that you see in here, so we have the head
tags, and the HTML tags and everything else there.

38
00:04:36,980 --> 00:04:42,800
And now here is where the BeautifulSoup
comes into play.

39
00:04:42,800 --> 00:04:49,820
So all the request does is it loaded
with the source code of a web page but

40
00:04:49,820 --> 00:04:56,270
in a very scrambled form as you see here.
Now if you want to make this beautiful

41
00:04:56,270 --> 00:05:02,120
and extract the elements, and the
text, and everything out of this source

42
00:05:02,120 --> 00:05:08,090
code, you want to use BeautifulSoup,
so all BeautifulSoup does is parsing

43
00:05:08,090 --> 00:05:14,600
this source code and giving you what you
want, so giving you the elements of the

44
00:05:14,600 --> 00:05:21,500
HTML tags your interesting about, so you
have already loaded this content and now

45
00:05:21,500 --> 00:05:28,820
what you want to do is maybe create
a variable and call it soup. And that would

46
00:05:28,820 --> 00:05:34,070
be equal to the beautiful soup and guess
what you want to pass here? Well that

47
00:05:34,070 --> 00:05:42,170
would be the content and maybe an another
argument so you want to specify the parser

48
00:05:42,170 --> 00:05:47,960
you want to use for parsing
these data, and that is normally the

49
00:05:47,960 --> 00:05:55,010
HTML parser, so this assault you want to
use almost always. If you don't specify

50
00:05:55,010 --> 00:06:00,590
this, you'll get a warning but still things
will work, so I normally pass it there.

51
00:06:00,590 --> 00:06:08,360
And once you've done that so I execute
that cell, if you know print soup

52
00:06:08,360 --> 00:06:17,990
dot prettify with empty brackets there,
you'll see the source code of the web

53
00:06:17,990 --> 00:06:26,090
page in an organized form, so BeautifulSoup
is trained to actually recognize

54
00:06:26,090 --> 00:06:32,020
these tags and then render them in
a visual way for the human eye.

55
00:06:32,020 --> 00:06:37,479
However this is just for demonstration.
Normally you will not have to actually

56
00:06:37,479 --> 00:06:46,120
use the prettify method a lot because
a better method to see this code as I

57
00:06:46,120 --> 00:06:50,620
already mentioned before is to… Let me
delete this cell. We don't need that,

58
00:06:50,620 --> 00:06:56,430
so a better way to see that source code is
to go to your web page and go to inspect

59
00:06:56,430 --> 00:07:05,880
and here you see a more, a better syntax
of the HTML code. So here you'll see that

60
00:07:05,880 --> 00:07:13,440
we have three divisions here with the cities
class, we have some more divisions here

61
00:07:13,440 --> 00:07:23,770
but this is what we're interested about.
So and the body is everything and if you

62
00:07:23,770 --> 00:07:30,520
expand one of these divisions, you'll see
that we have an h2 tag, so a heading tag

63
00:07:30,520 --> 00:07:38,830
and also a paragraph tag so P tag and h2
tags, and also the other division which

64
00:07:38,830 --> 00:07:44,050
is this one here has this h2 tag and the
paragraph tag, and Tokyo also has the

65
00:07:44,050 --> 00:07:54,280
same thing. So our duty now is to extract
the names of these elements so that

66
00:07:54,280 --> 00:08:02,020
would be the h2, the text of the h2
tags inside the cities tags.

67
00:08:02,020 --> 00:08:07,870
So naturally you'll start thinking about
iterating through these boxes which are

68
00:08:07,870 --> 00:08:13,449
actually divisions. So you want to go
through here, here, and here and extract

69
00:08:13,449 --> 00:08:22,770
what you want to extract. So we go back
to the code and what you want to do is

70
00:08:22,770 --> 00:08:34,570
perform a method called find all, and
what you want to find is divs, so divs

71
00:08:34,570 --> 00:08:40,449
but there may be lots of divs in the
web page so for instance we have two

72
00:08:40,449 --> 00:08:45,850
more divs here and we don't want these
to be found, we only one these three.

73
00:08:45,850 --> 00:08:53,019
But these three as you see they
have a common class attribute

74
00:08:53,019 --> 00:08:58,300
which is equal to cities.
So we want to make use of that and we

75
00:08:58,300 --> 00:09:11,410
pass here a dictionary which would be class
equals to cities. Okay, let me create

76
00:09:11,410 --> 00:09:22,660
a variable here and call it all and
execute it. Now if you print all, you'll

77
00:09:22,660 --> 00:09:28,870
see that the divisions have been
extracted from the source code so from

78
00:09:28,870 --> 00:09:35,339
the soup which was the entire source
code and I'd like you to actually see

79
00:09:35,339 --> 00:09:41,110
closely here. You can see that the
first division is divided by comma here

80
00:09:41,110 --> 00:09:48,190
and then the second division starts.
So for Paris, Paris is the second and it

81
00:09:48,190 --> 00:09:53,649
ends here, and then Tokyo starts here so
we've got a list with three elements, one

82
00:09:53,649 --> 00:10:00,670
for each division. Now if you want to
find only the first element with this

83
00:10:00,670 --> 00:10:11,319
class attribute, so cities you'd want to
use the find method. All, and in this case

84
00:10:11,319 --> 00:10:16,930
you don't get a list, but you get the
code for the division, for the first

85
00:10:16,930 --> 00:10:27,250
division only which happens to be a tag
element of BeautifulSoup, so it's not

86
00:10:27,250 --> 00:10:31,899
a plain string but it's a special,
let's say a special BeautifulSoup

87
00:10:31,899 --> 00:10:38,920
string so that BeautifulSoup knows its
structure, so it knows what are elements,

88
00:10:38,920 --> 00:10:42,759
where the tags are, and where the text is
and so on, so that BeautifulSoup is

89
00:10:42,759 --> 00:10:50,050
able to give you the information
that you are looking for, so all again.

90
00:10:50,050 --> 00:10:55,990
So you extract the first element.
Now an alternative way to extract the

91
00:10:55,990 --> 00:11:03,870
first element is logically,
so you have all elements here, is to use

92
00:11:03,870 --> 00:11:11,940
the list indexing. So this object that
I just showed you, the tag object of

93
00:11:11,940 --> 00:11:19,770
BeautifulSoup supports indexing.
So you execute that and in this

94
00:11:19,770 --> 00:11:26,940
case as you see you extracted the first
item of the tag object or you could do

95
00:11:26,940 --> 00:11:33,660
it like this. So you grab all of them so
here you have all of them and 0 is the

96
00:11:33,660 --> 00:11:40,520
first one, you get the idea.
Ok, but what if you want only the h2 tags

97
00:11:40,520 --> 00:11:48,510
from this div class? Well in that case
what you'd want to do this refer to the

98
00:11:48,510 --> 00:11:55,830
all objects and then apply the find all
method again and this time you'd want to

99
00:11:55,830 --> 00:12:01,470
get the h2 element. And in this case you
don't have a class attribute so you'll

100
00:12:01,470 --> 00:12:07,320
have to leave it like that, and you get
an error because what I did here is I

101
00:12:07,320 --> 00:12:17,550
didn't point to this division but
I pointed to actually the list containing

102
00:12:17,550 --> 00:12:24,450
all these divisions, so Python is trying
to get the h2 but this result set method

103
00:12:24,450 --> 00:12:31,620
doesn't have this h2 element, so what you
want to do is you want to point to the

104
00:12:31,620 --> 00:12:37,709
first element of the list and that
gives you the h2 element with the tags

105
00:12:37,709 --> 00:12:46,110
and the text which is like a list so you
want to perform a zero indexing there.

106
00:12:46,110 --> 00:12:54,899
And if you want London only you apply
text and you get London. So this is what we

107
00:12:54,899 --> 00:13:01,170
wanted, right? To extract the cities.
So we extracted London. Now how about

108
00:13:01,170 --> 00:13:08,070
extracting Paris and Tokyo? Well as
you might guess we need to use a for

109
00:13:08,070 --> 00:13:12,089
loop, but first let me summarize
what we did here.

110
00:13:12,089 --> 00:13:18,809
So we loaded the content up here which
this this one here, and then we loaded this

111
00:13:18,809 --> 00:13:23,550
content in the BeautifulSoup method.
And BeautifulSoup makes

112
00:13:23,550 --> 00:13:29,730
this soup beautiful so that it
recognizes these tags and so what we did

113
00:13:29,730 --> 00:13:37,620
then is we found, we extracted from this
content, we extracted all the division

114
00:13:37,620 --> 00:13:42,930
elements, so together with the tags and
the attributes, and the text inside them.

115
00:13:42,930 --> 00:13:48,110
So everything inside these divisions
with a class equals to cities.

116
00:13:48,110 --> 00:13:57,749
Then we can perform for each of
these elements of this list we can

117
00:13:57,749 --> 00:14:05,220
perform again a find all method, so we can
find subtags of these division tags.

118
00:14:05,220 --> 00:14:10,800
In this case we found the h2 tags and
then we grab the first item of the list

119
00:14:10,800 --> 00:14:15,779
which in this case happens to be a list
with only one item, so each of these

120
00:14:15,779 --> 00:14:21,569
divisions have one h2 tags or
alternatively you could just use find

121
00:14:21,569 --> 00:14:28,110
here and without using this indexing but
this is a general method and then we

122
00:14:28,110 --> 00:14:34,589
apply the text attribute there so to
extract the text out of this element.

123
00:14:34,589 --> 00:14:42,809
So we got London. Now we need to do the same,
but not in this case by iterating so for

124
00:14:42,809 --> 00:14:55,559
let's say item in all, you want to print
out, so item here is this one here,

125
00:14:55,559 --> 00:15:03,199
so this is this would be the first item,
so you want to print out the item dot find

126
00:15:03,199 --> 00:15:11,579
all and you want to find the h2 tags
from this first item for example.

127
00:15:11,579 --> 00:15:19,290
So h2 tags and then you need to
apply this zero indexing there and you

128
00:15:19,290 --> 00:15:26,050
want to grab the text from this and
that's it. Here are the data.

129
00:15:26,050 --> 00:15:33,950
Alternatively you could just pass P here
and you get the paragraphs, so these ones

130
00:15:33,950 --> 00:15:41,740
here, the text. And so that's the idea of
loading web pages in Python and

131
00:15:41,740 --> 00:15:48,400
parsing them with BeautifulSoup
and extracting text out of the web page.

132
00:15:48,400 --> 00:15:55,160
So sorry if I was a bit repetitive
in explaining this stuff, but I really

133
00:15:55,160 --> 00:16:00,860
want to make sure you understand the
core concepts, on the other hand if you

134
00:16:00,860 --> 00:16:06,980
found these very basic, I would say let's
move on to the next lectures where we

135
00:16:06,980 --> 00:16:12,290
will be extracting some information
from a more advanced website and

136
00:16:12,290 --> 00:16:17,690
we'll be extracting links and not only
text, so that's a real world program and

137
00:16:17,690 --> 00:16:22,000
a very interesting one.
So I'll talk to you later.

