1
00:00:00,060 --> 00:00:06,870
Great, so we got the script working for
one page, so for the first page, but we

2
00:00:06,870 --> 00:00:13,290
may have multiple pages there, so like for
Rock Springs we have three pages with 28

3
00:00:13,290 --> 00:00:19,920
listings, so 28 properties listed there
for sale or for rent. Actually they are for

4
00:00:19,920 --> 00:00:27,060
sales, so homes for sale. So this is the
second page with 10 more properties and

5
00:00:27,060 --> 00:00:34,079
then you have the last page there which
should have eight properties, so we have

6
00:00:34,079 --> 00:00:40,260
28 in total. Now this is a small town so it
doesn't have many listings there,

7
00:00:40,260 --> 00:00:44,489
but with big cities you'll have lots of
pages and the script that we'll be

8
00:00:44,489 --> 00:00:49,770
building to extract properties, property
data from these pages will work with any

9
00:00:49,770 --> 00:00:56,129
number of pages. How do we go about
grabbing data from all the pages?

10
00:00:56,129 --> 00:01:01,800
Well there's no magic there, but there is
a simple trick. You know if you search

11
00:01:01,800 --> 00:01:08,430
again for Rock Springs you'll see an
initial URL there which we loaded in

12
00:01:08,430 --> 00:01:16,490
Python. I think it was without this
k equals to 1 so it was like this only.

13
00:01:16,490 --> 00:01:23,400
So I'm not sure what the k meant there, but
anyway maybe it was because I searched

14
00:01:23,400 --> 00:01:30,079
again for Rock Springs. So this is the
first page. Now the logic behind

15
00:01:30,079 --> 00:01:36,600
extracting the data from the other pages
is what we need to do is we need to load

16
00:01:36,600 --> 00:01:44,040
the other pages in Python as well using
requests again, so just as we did here,

17
00:01:44,040 --> 00:01:47,970
we loaded the first page. Now we need to
figure out a way to load the other pages.

18
00:01:47,970 --> 00:01:54,689
We can either go manually through
each of the 300 pages, three pages in

19
00:01:54,689 --> 00:02:02,520
this case and copy the URL, or we can try
to figure out a way, so to figure out the

20
00:02:02,520 --> 00:02:07,200
rule if there is a rule that the URL
changes when you switch through the

21
00:02:07,200 --> 00:02:15,020
pages and it probably always be a rule
there, so we need to find that out.

22
00:02:20,810 --> 00:02:26,569
So I went to the second page and as you
see now the URL changed with some

23
00:02:26,569 --> 00:02:33,830
extension there so T equals to zero and
S equals to zero. If you go to the third

24
00:02:33,830 --> 00:02:45,489
page, you see that the URL before 20
didn't change, but S was changed to 20.

25
00:02:45,489 --> 00:02:55,180
Now if you go to the first page again
you'll see that S is equal to zero.

26
00:02:55,730 --> 00:03:02,459
So when you search for Rock Springs
directly you get the base URL there

27
00:03:02,459 --> 00:03:08,580
without this extension, and then when you
go to the first page manually by

28
00:03:08,580 --> 00:03:13,530
clicking the number one here, you'll be
included in the rule that changes the

29
00:03:13,530 --> 00:03:19,560
URL when you move throughout the pages.
That means we can use this URL to grab

30
00:03:19,560 --> 00:03:27,480
the first page then we can loop through
the URLs and change this value to 10

31
00:03:27,480 --> 00:03:37,380
as we iterate, so increase by 10, so from
0, 10, 20 and if we have more pages that

32
00:03:37,380 --> 00:03:47,489
would be 30 and so on. So I hope you get
the idea. Let's now go ahead and let's

33
00:03:47,489 --> 00:03:55,590
say, so let's store a base URL there
which would be something like this without

34
00:03:55,590 --> 00:04:00,230
the number, so this is the stat URL.
And now before we request the source

35
00:04:07,470 --> 00:04:14,670
code with requests, let's first try to
build URLs and print them out, so we're

36
00:04:14,670 --> 00:04:21,120
using a print statement as always to try
things out so let's say for page in

37
00:04:21,120 --> 00:04:30,930
range, well the range would start from 0
and up to, we had 3 pages there and last

38
00:04:30,930 --> 00:04:39,870
page had an S of 20. So maybe put
30 there and then iterate with a step of 10.

39
00:04:39,870 --> 00:04:46,260
So at each iteration we will increase
by 10. For now don't worry about the

40
00:04:46,260 --> 00:04:52,170
number of pages. For now we are just
putting a value manually, so we know that

41
00:04:52,170 --> 00:04:57,090
this has three pages, so I'm putting 30
there, but later on we will change this

42
00:04:57,090 --> 00:05:02,610
to actually a variable that grabs a
number of pages from the website, so if

43
00:05:02,610 --> 00:05:07,320
instead of Rock Springs you'd have New
York and you'd get probably 200 pages,

44
00:05:07,320 --> 00:05:12,660
so you want to grab the number of the
last page, so 200 and then pass something

45
00:05:12,660 --> 00:05:19,800
like 200 times 20 here, so you get 2 000
and you get a loop that goes to 0, 10, and

46
00:05:19,800 --> 00:05:30,180
to 2 000, so we'll do that later.
For now print base URL + well this page

47
00:05:30,180 --> 00:05:36,030
variable would get an integer, so we need
to convert that integer to string using

48
00:05:36,030 --> 00:05:41,180
the string method, and then page goes
inside that.

49
00:05:44,220 --> 00:05:57,480
So let's try this out. And yeah, these are
the three links. Let me open one of them, here we

50
00:05:57,480 --> 00:06:09,050
go, and this is the second set of results.
So it's page 2, great, let me close this.

51
00:06:09,050 --> 00:06:14,220
But of course printing is not what we're
interested about so we don't want to print

52
00:06:14,220 --> 00:06:20,160
the URLs, we want to get the source
code of them, so what you want to do is

53
00:06:20,160 --> 00:06:27,180
R equal to requests, so you're inside
the loop and later we will merge this

54
00:06:27,180 --> 00:06:37,020
loop with the other code so let's focus
on this for now. Get base URL plus the

55
00:06:37,020 --> 00:06:41,240
string of page number.
And just after that you want to grab

56
00:06:44,349 --> 00:06:56,769
the content of requests and also make
the soup, so Beautiful Soup, C and HTML

57
00:06:56,769 --> 00:07:06,849
parser, and for now let's try to print out
the prettified version of the soup

58
00:07:06,849 --> 00:07:16,989
object. Great! And that took a while but
here is the result. So in the first

59
00:07:16,989 --> 00:07:21,879
iteration we got the URL printed out by
this function, by the print function and

60
00:07:21,879 --> 00:07:29,169
then we get the prettified version of
the source code for first URL which is

61
00:07:29,169 --> 00:07:39,099
quite long, so anyway you get the idea.
So that seems to be working, but what we

62
00:07:39,099 --> 00:07:45,039
need from all this source code, so we
have a source code of all three page,

63
00:07:45,039 --> 00:07:50,829
so we don't need to print all the source
code, but we only need to extract the

64
00:07:50,829 --> 00:07:57,759
property row divisions so just as we did
here, so we got the soup for the first page

65
00:07:57,759 --> 00:08:07,469
and then we created an all variable
there, so we do the same here, but this

66
00:08:07,469 --> 00:08:13,629
time it should be indented because it's
a inside a loop. So all and property row

67
00:08:13,629 --> 00:08:23,079
class division. That looks good.
So if you print it just to check all and

68
00:08:23,079 --> 00:08:28,569
here is the result, so the first URL
printed out, and then you have these

69
00:08:28,569 --> 00:08:37,689
divisions and you have quite a lot of
data here as well. So now what you want

70
00:08:37,689 --> 00:08:44,050
to do is you want to grab all this code
that you just built there and maybe delete

71
00:08:44,050 --> 00:08:53,430
this and delete the entire cell and
you want to put that code up here.

72
00:08:53,769 --> 00:09:01,790
Great, and now next thing is you know you
did this for one page, and now you're

73
00:09:01,790 --> 00:09:06,879
looping through all the pages, so what
you want to do is you want to indent

74
00:09:06,879 --> 00:09:13,600
all this block on the right so that this
block becomes part of the for loop there.

75
00:09:13,600 --> 00:09:21,199
So control and the closing square brackets
to indent the selected text, so now this

76
00:09:21,199 --> 00:09:29,600
loop here wheeler on as many times as
there are pages there. So what can we

77
00:09:29,600 --> 00:09:39,259
do now? Well we can try out the code so
let me execute this. Shift, enter and we

78
00:09:39,259 --> 00:09:45,769
have to wait a bit yeah. So that looks
like it's working and so at this point

79
00:09:45,769 --> 00:09:52,339
we should have the L list there with all
the data and here we convert that L list to

80
00:09:52,339 --> 00:09:58,910
a data frame so let's execute this and
then here we execute the data frame.

81
00:09:58,910 --> 00:10:10,639
So let's see now. Yep!
And so we have 24, 25 rows because we

82
00:10:10,639 --> 00:10:13,569
start from zero,
but we had 28 listings actually so that

83
00:10:18,050 --> 00:10:25,430
means this last three properties have
not been listed here because something

84
00:10:25,430 --> 00:10:32,689
happened there, so we are in page 3 here
and the last successful one was Wendt

85
00:10:32,689 --> 00:10:39,410
Avenue, so you look for Wendt Avenue. Here is
Wendt Avenue, so it has the address and

86
00:10:39,410 --> 00:10:45,860
the locality, so Rock Springs and then
the one after that it has only the

87
00:10:45,860 --> 00:10:52,100
locality there, so in that case what you
could do is you could try to handle an

88
00:10:52,100 --> 00:11:02,829
exception here in the code, so for this
one here we could say try to do that,

89
00:11:02,829 --> 00:11:16,490
and accept d locality equals to none.
Now I'll execute this again. Let's hope we

90
00:11:16,490 --> 00:11:23,870
don't get an error there. Yeah, that was
successfully executed. Execute that and the

91
00:11:23,870 --> 00:11:38,360
data frame and let's see. Yeah, now we seem to
have 28 rows there so that seems to be

92
00:11:38,360 --> 00:11:45,199
working, and the last thing you want
to do is you know instead of passing

93
00:11:45,199 --> 00:11:49,550
this range manual here, we could add
actually a variable that detects the

94
00:11:49,550 --> 00:12:00,860
page number, and all we have to do is we
go to our website, let me remove this, so

95
00:12:00,860 --> 00:12:06,470
you want to figure out where this last
number is being rendered so that is the

96
00:12:06,470 --> 00:12:11,120
indicator of the number of pages, so we
have three pages in this case. Inspect

97
00:12:11,120 --> 00:12:19,509
that, so that brings us to this point.
Actually in this case it says, page,

98
00:12:19,509 --> 00:12:27,620
current page, so let's go to the
first page and let me close this.

99
00:12:27,620 --> 00:12:32,300
We are supposed to be the first page
when we run the script for the first time

100
00:12:32,300 --> 00:12:37,340
and then in that first page, so we want to
grab the content of the first page, and

101
00:12:37,340 --> 00:12:45,410
then we want to know the element of that
last number there, so we see that that's

102
00:12:45,410 --> 00:12:53,210
the A element with class page, and here
is the text, so three, and that should be

103
00:12:53,210 --> 00:12:59,270
actually the last element with a class
of page so we'll have quite a few

104
00:12:59,270 --> 00:13:06,260
elements with class of page there, and
this should be last one. Yeah, so that

105
00:13:06,260 --> 00:13:11,120
should be the last one so keep that in mind.
And then we go here and what we can do here

106
00:13:11,120 --> 00:13:18,020
now is, or let's grab that page number in
here, so page, let's call that page number.

107
00:13:18,020 --> 00:13:25,910
That would be equal to soup, so soup will
contain the HTML code of the first page,

108
00:13:25,910 --> 00:13:32,150
so this one here which is modern Ajax
request, so it's a simple one of this

109
00:13:32,150 --> 00:13:50,450
page and so we want to find all
A elements with a class of page, so if I

110
00:13:50,450 --> 00:13:58,670
print it out just to see what we have so
far, page number. See we've got a few

111
00:13:58,670 --> 00:14:07,250
elements there, but the last one is the
lost page number so three. That's what we

112
00:14:07,250 --> 00:14:13,370
need, so we need to grab the last
item of the list which means -1, an index

113
00:14:13,370 --> 00:14:23,240
of -1 and text. Yep, we got three here and
so we need to transform this number now

114
00:14:23,240 --> 00:14:30,890
into 30 and the way you do that is just
page number times 10.

115
00:14:30,890 --> 00:14:33,650
So that would give you 30.
Well, almost because this page number is

116
00:14:33,650 --> 00:14:38,260
actually a string here. You know type,
another bracket, so is type string and

117
00:14:46,600 --> 00:14:54,200
that means you need to add integer here.
So you want to convert this string to an

118
00:14:54,200 --> 00:15:00,350
integer and then multiplied by 10 so
that you get you get 30. So that should

119
00:15:00,350 --> 00:15:11,240
do it now, and everything should work I
believe, so let me execute this. The URL

120
00:15:11,240 --> 00:15:15,290
are being printed out and the script
finished so when you don't have th busy

121
00:15:15,290 --> 00:15:19,850
text before the title of your file here
that means the script is running so now

122
00:15:19,850 --> 00:15:26,330
we don't have that, so the script
finished. Execute that and this is

123
00:15:26,330 --> 00:15:39,460
our results, and export it to a CSV file,
and let's check, so output dot CSV.

124
00:15:40,570 --> 00:15:51,860
And here is the data. So we don't have any
repetitions here. No, everything looks

125
00:15:51,860 --> 00:16:01,790
unique. And that closes this section of
web scrapping, so if you want to find

126
00:16:01,790 --> 00:16:07,280
other localities, so I tried this for
Rock Springs, but you can try it with

127
00:16:07,280 --> 00:16:12,820
any other localities, so all you have to
do is change the URL, or you could even

128
00:16:12,820 --> 00:16:17,210
implement some user input so the user
enters the name, and then you construct the

129
00:16:17,210 --> 00:16:27,760
URL, and then you pass the URL in here.
So we don't need this. Clean this up as well.

130
00:16:27,760 --> 00:16:33,050
And if you want you can merge this cell
with this one, so you can execute the

131
00:16:33,050 --> 00:16:39,980
entire script in one cell and you can
also save this, so download it as

132
00:16:39,980 --> 00:16:46,520
a Python file and you'll get all the source
code, so the Python script and plus

133
00:16:46,520 --> 00:16:53,389
these lines, but with a comment hashtag.
And you can run it as Python file

134
00:16:53,389 --> 00:16:58,820
the entire script. So I hope you enjoyed
this! I know this was quite a bit to

135
00:16:58,820 --> 00:17:05,600
digest, and if you have questions please
ask them! I'd be happy to help you! So I

136
00:17:05,600 --> 00:17:10,220
hope you learned a lot from this, and we'll
move on through the next sections, so we

137
00:17:10,220 --> 00:17:15,000
still have very interesting things to do.
And I'll see you later!

