1
1

00:00:01,990  -->  00:00:04,610
<v Jose>In this video we are going to use BeautifulSoup</v>
2

2

00:00:04,610  -->  00:00:09,610
to get the price that is present in this website.
3

3

00:00:09,890  -->  00:00:13,194
What we have to do is, of course, import BeautifulSoup.
4

4

00:00:13,194  -->  00:00:16,830
So, we're going to say from bs4 import BeautifulSoup,
5

5

00:00:17,860  -->  00:00:22,690
and then, we're going to define a couple of constants.
6

6

00:00:22,690  -->  00:00:27,260
So, I'm gonna cut that out and make it into a constant,
7

7

00:00:27,260  -->  00:00:30,470
and then, we're gonna create a constant for the tag name
8

8

00:00:30,470  -->  00:00:34,850
which was a paragraph tag, a P tag, in the website.
9

9

00:00:34,850  -->  00:00:37,010
And then, we're going to have a query,
10

10

00:00:37,010  -->  00:00:41,533
which is, class is price, and price, dash, dash, large.
11

11

00:00:42,400  -->  00:00:46,950
So, together we have the paragraph tag, and the class name.
12

12

00:00:46,950  -->  00:00:51,110
This query, which is a dictionary that has class as the key,
13

13

00:00:51,110  -->  00:00:54,790
and the value, as the value, is what BeautifulSoup
14

14

00:00:54,790  -->  00:00:58,500
is going to use to identify a specific thing in the page.
15

15

00:00:58,500  -->  00:01:01,110
So, that's why we define that in that format,
16

16

00:01:01,110  -->  00:01:03,500
this is what BeautifulSoup expects.
17

17

00:01:03,500  -->  00:01:06,650
So, now we're gonna do requests get of the URL
18

18

00:01:06,650  -->  00:01:08,690
that we've defined up here.
19

19

00:01:08,690  -->  00:01:11,140
Then we're going to say that the content of the page
20

20

00:01:11,140  -->  00:01:13,633
is the response, dot, content.
21

21

00:01:14,870  -->  00:01:18,430
Then we're going to tell BeautifulSoup to understand,
22

22

00:01:18,430  -->  00:01:21,670
or to read the HTML that we've received.
23

23

00:01:21,670  -->  00:01:24,553
So, we're gonna say soup is BeautifulSoup of the content,
24

24

00:01:25,413  -->  00:01:28,489
and we also have to tell it that this content is HTML.
25

25

00:01:28,489  -->  00:01:30,578
So, the way we do that is we say,
26

26

00:01:30,578  -->  00:01:33,410
"BeautifulSoup, here you go with some HTML,"
27

27

00:01:33,410  -->  00:01:38,410
and by the way, use the HTML parser for that content.
28

28

00:01:38,840  -->  00:01:42,330
So, the way we do that is we give it a string with
29

29

00:01:42,330  -->  00:01:47,330
HTML, dot, parser in it to the BeautifulSoup constructor.
30

30

00:01:47,340  -->  00:01:49,290
Then we can find the specific element
31

31

00:01:49,290  -->  00:01:51,100
that matches our search.
32

32

00:01:51,100  -->  00:01:53,498
So, we're gonna say soup, dot, find,
33

33

00:01:53,498  -->  00:01:58,498
and then we give it the tag name, and we give it the query.
34

34

00:01:58,640  -->  00:02:02,730
That is going to go ahead and find the element that has
35

35

00:02:02,730  -->  00:02:07,230
the P tag, and this class here.
36

36

00:02:07,230  -->  00:02:10,410
Then, we're going to find the price
37

37

00:02:10,410  -->  00:02:12,490
that is contained within that element.
38

38

00:02:12,490  -->  00:02:16,320
So, that was some text, and BeautifulSoup allows us
39

39

00:02:16,320  -->  00:02:18,993
to get the text of an element.
40

40

00:02:19,860  -->  00:02:22,180
So, now that we've got this, we're gonna print it out,
41

41

00:02:22,180  -->  00:02:23,780
and see what happens.
42

42

00:02:23,780  -->  00:02:25,580
So, I'm gonna press play once again,
43

43

00:02:26,500  -->  00:02:27,333
and there you have it.
44

44

00:02:27,333  -->  00:02:30,760
You have the 1,469.00.
45

45

00:02:30,760  -->  00:02:33,220
Very important, this is clearly a string.
46

46

00:02:33,220  -->  00:02:36,160
It is not a number because we've got the pound sign,
47

47

00:02:36,160  -->  00:02:39,710
and we've got the comma, and in Python those things
48

48

00:02:39,710  -->  00:02:41,300
are not present in numbers.
49

49

00:02:41,300  -->  00:02:43,290
Something else to remember, is that this number
50

50

00:02:43,290  -->  00:02:45,933
actually has a bunch of white space in it.
51

51

00:02:45,933  -->  00:02:50,667
It's not just the number, it is also a bunch of spaces.
52

52

00:02:51,980  -->  00:02:54,500
We can get rid of any spaces that are behind,
53

53

00:02:54,500  -->  00:02:58,690
or in front of a number, by doing dot, strip, on a string.
54

54

00:02:58,690  -->  00:03:01,360
So it's element, dot, text is a string,
55

55

00:03:01,360  -->  00:03:03,050
element, dot, text, dot, strip
56

56

00:03:03,050  -->  00:03:04,610
is going to remove any white space.
57

57

00:03:04,610  -->  00:03:06,630
So, if we run this again, you can see
58

58

00:03:06,630  -->  00:03:09,730
that now this is slightly different.
59

59

00:03:09,730  -->  00:03:12,290
Now that we have this price, we are almost there.
60

60

00:03:12,290  -->  00:03:15,890
We can almost get the price from the website.
61

61

00:03:15,890  -->  00:03:17,500
All that we need to do next,
62

62

00:03:17,500  -->  00:03:19,806
is use some regular expressions,
63

63

00:03:19,806  -->  00:03:24,806
in order to extract a number from this string.
64

64

00:03:25,450  -->  00:03:27,480
And we're going to do that in the next video.
65

65

00:03:27,480  -->  00:03:28,630
So, I'll see you there.
