1
00:00:00,030 --> 00:00:05,700
So how does web scrapping work anyway?
How is Python able to grab the

2
00:00:05,700 --> 00:00:12,530
information from the webpage and store it as
text so that you can analyze it in your

3
00:00:12,530 --> 00:00:18,840
convenience?
Well this text that you see here, so this

4
00:00:18,840 --> 00:00:26,580
is an example website and actually this
domain is reserved for examples so but

5
00:00:26,580 --> 00:00:31,730
it's a normal web page built with HTML,
and CSS, and other word development tools.

6
00:00:31,730 --> 00:00:37,530
So HTML is what renders the elements,
the text, and everything on the browser.

7
00:00:37,530 --> 00:00:46,230
Now luckily you can see the code of
every web page by going to view page

8
00:00:46,230 --> 00:00:52,770
source as I did here, so this is the code,
the HTML code as you see it opens

9
00:00:52,770 --> 00:00:59,390
here with HTML tags and it closes here.
And so this has the title of the web page

10
00:00:59,390 --> 00:01:06,119
which is this one here example domain
and so on. And most importantly for you

11
00:01:06,119 --> 00:01:13,470
to know is that the HTML script is made
of HTML elements, so this line here is an

12
00:01:13,470 --> 00:01:21,060
HTML element so it's a meta element and
then we have division elements which is

13
00:01:21,060 --> 00:01:27,689
this one here. We have paragraph elements,
these here and these are called tags so

14
00:01:27,689 --> 00:01:33,270
the div tag and the division
closing tag, so opening division

15
00:01:33,270 --> 00:01:39,689
tag, closing division tag and then we
have the body tag, so here is where the

16
00:01:39,689 --> 00:01:46,380
visible part of the HTML page, of a web
page is placed so everything you place

17
00:01:46,380 --> 00:01:52,740
inside the body tags you'll see there so
if you have example domain here, this

18
00:01:52,740 --> 00:01:58,079
domain established to be usable etc.,
here is where you see that text.

19
00:01:58,079 --> 00:02:07,110
This domain is established to be used for etc.
So the HTML elements are key for web

20
00:02:07,110 --> 00:02:12,510
scrapping, so basically what you'll do
is that

21
00:02:12,510 --> 00:02:21,599
let's say you want to extract the text
of h1 tags out of all the divs, all the

22
00:02:21,599 --> 00:02:28,980
divisions, so what you tell Python to do
is you tell Python to go to all the h1

23
00:02:28,980 --> 00:02:35,519
tags and extract the text of these tags
and Python will do that, but first of all

24
00:02:35,519 --> 00:02:40,709
though you need to load this entire
script in Python and a way to do that is

25
00:02:40,709 --> 00:02:47,670
by using the request library, so request
allows you to give Python a URL so like

26
00:02:47,670 --> 00:02:53,280
example.com and Python will grab all
this text, and then once you have this

27
00:02:53,280 --> 00:03:01,730
text, you'll use the beautiful soup library
to extract all the elements from the tags.

28
00:03:01,730 --> 00:03:08,220
So for instance the example, the text
inside h1 tag, and then

29
00:03:08,220 --> 00:03:12,750
you can store the extracted text in
variables, or in Python dictionaries,

30
00:03:12,750 --> 00:03:16,440
or Pandas dataframes
and wherever you find it's useful for

31
00:03:16,440 --> 00:03:28,079
your needs. So that's the concept and you
can see the same source code as you

32
00:03:28,079 --> 00:03:34,620
might already know from the inspect, so
right-click inspect and you see this.

33
00:03:34,620 --> 00:03:40,560
So here is like this source code but you see
it's more organized so for instance if

34
00:03:40,560 --> 00:03:45,720
you're opening your mouse over the body
tags the corresponding element in the

35
00:03:45,720 --> 00:03:52,380
webpage will be highlighted as you see
it, so if you expand this now you are on

36
00:03:52,380 --> 00:04:00,060
the h1 tags and so on. So this allows you
to actually see the names of all the

37
00:04:00,060 --> 00:04:04,730
tags for the elements you want to
extract, so we'll use the inspect

38
00:04:04,730 --> 00:04:11,790
window for understanding the source code
of our web pages. So that's about the

39
00:04:11,790 --> 00:04:16,139
concept of web scrubbing and I'll see
you in the next lecture where I'll load

40
00:04:16,139 --> 00:04:21,599
a web page in Python and then we will
extract some simple data from the web

41
00:04:21,599 --> 00:04:26,400
page, so just to get you started with the
requests library

42
00:04:26,400 --> 00:04:30,000
and the BeautifulSoup library as well.
So let's move on.

