WEBVTT 1 00:00:00.210 --> 00:00:01.430 Hi and welcome back, 2 00:00:01.430 --> 00:00:04.940 let's review our quotes project in this video. 3 00:00:06.400 --> 00:00:08.680 Instead of starting from the quote parser 4 00:00:08.680 --> 00:00:11.570 let's start from the top of the project now. 5 00:00:11.570 --> 00:00:14.890 The first thing we do, is aside from importing things 6 00:00:14.890 --> 00:00:16.810 we get the page content, 7 00:00:16.810 --> 00:00:19.010 this is the entire HTML content 8 00:00:19.010 --> 00:00:20.150 of the entire QuotesPage 9 00:00:22.150 --> 00:00:26.520 We then give that content to our QuotesPage constructor, 10 00:00:26.520 --> 00:00:28.870 the in its method, so let's go over there, 11 00:00:30.140 --> 00:00:34.680 that then initialises this BeautifulSoup object, 12 00:00:34.680 --> 00:00:36.240 using the HTML parser 13 00:00:36.240 --> 00:00:39.180 so this lets us search within this page. 14 00:00:40.930 --> 00:00:45.460 The quotes property of that finds a locator 15 00:00:45.460 --> 00:00:47.740 which is this QuotesPageLocators dot Quote 16 00:00:48.700 --> 00:00:53.700 which is just div dot quote, okay. 17 00:00:55.400 --> 00:00:58.700 Then it selects all the instances 18 00:00:58.700 --> 00:01:02.700 of that selector so every div that has a quote class 19 00:01:02.700 --> 00:01:04.910 will be selected by this, 20 00:01:04.910 --> 00:01:07.710 and if we go back to Chrome we can see 21 00:01:07.710 --> 00:01:10.000 that the divs that have a class quote 22 00:01:10.900 --> 00:01:13.820 are all the individual quotes. 23 00:01:16.480 --> 00:01:19.400 And it's gonna return a QuoteParser 24 00:01:19.400 --> 00:01:21.920 given each of those tags 25 00:01:22.940 --> 00:01:25.770 for e in the tags that we found. 26 00:01:25.770 --> 00:01:27.660 Remember the tags that we found there 27 00:01:27.660 --> 00:01:31.300 are not strings, they are BeautifulSoup tags, 28 00:01:32.200 --> 00:01:34.440 so if we go into the quote parser, 29 00:01:34.440 --> 00:01:37.290 this parent will be a BeautifulSoup tag 30 00:01:37.290 --> 00:01:39.080 which is that div. 31 00:01:40.620 --> 00:01:43.470 In BeautifulSoup, everything is a tag, 32 00:01:43.470 --> 00:01:46.000 so when you load the entire HTML page, 33 00:01:46.000 --> 00:01:47.720 that's a tag that you can search on, 34 00:01:47.720 --> 00:01:50.040 that's an HTML tag. 35 00:01:50.040 --> 00:01:51.490 When you load one of these parents 36 00:01:51.490 --> 00:01:52.890 when you find one of these divs, 37 00:01:52.890 --> 00:01:54.510 that's just another tag, in this case, 38 00:01:54.510 --> 00:01:58.260 it's a div tag that you can search on as well. 39 00:01:58.260 --> 00:02:01.430 So now this quote parser has one of these divs 40 00:02:02.490 --> 00:02:06.720 and in their properties here we are searching 41 00:02:06.720 --> 00:02:10.300 under this parent only and we're searching for 42 00:02:10.300 --> 00:02:14.700 the span dot text, the small dot author, 43 00:02:14.700 --> 00:02:16.330 and the different tags. 44 00:02:17.340 --> 00:02:20.680 In addition, this quote parser also has a repr method 45 00:02:20.680 --> 00:02:22.850 that allows us to print it out nicely, 46 00:02:22.850 --> 00:02:24.730 but of course we can choose to print out 47 00:02:24.730 --> 00:02:27.080 specific properties if we choose. 48 00:02:29.720 --> 00:02:32.200 So hopefully this makes sense, 49 00:02:32.200 --> 00:02:34.080 the way this app is structured 50 00:02:34.080 --> 00:02:37.340 is to separate concerns, 51 00:02:37.340 --> 00:02:41.040 separate the locators away from your logic, 52 00:02:41.040 --> 00:02:45.750 and also separate the page from its components. 53 00:02:45.750 --> 00:02:48.280 In this case, we're separating the quotes page 54 00:02:48.280 --> 00:02:50.240 from an individual quote. 55 00:02:50.240 --> 00:02:52.530 Even though they're both sort of similar, 56 00:02:52.530 --> 00:02:54.550 I don't like calling a quote a page 57 00:02:54.550 --> 00:02:57.980 because really it's a component within a page. 58 00:02:57.980 --> 00:03:00.840 If you separate your scrapers in this way, 59 00:03:00.840 --> 00:03:02.960 I am confident you're going to be able 60 00:03:02.960 --> 00:03:07.360 to do anything you want with your scraping project. 61 00:03:07.360 --> 00:03:08.660 Just because doing this is going 62 00:03:08.660 --> 00:03:09.830 to make them so much simpler 63 00:03:09.830 --> 00:03:12.690 to work with and as soon as you want to add a new page, 64 00:03:12.690 --> 00:03:14.934 all you have to do is add a new file here 65 00:03:14.934 --> 00:03:17.270 and a new set of locators. 66 00:03:17.270 --> 00:03:19.230 If your locators change for any page, 67 00:03:19.230 --> 00:03:21.080 you just go there and modify them, 68 00:03:21.080 --> 00:03:24.070 and if you want to add a new component that you want 69 00:03:24.070 --> 00:03:26.280 to get data out of from a page, 70 00:03:26.280 --> 00:03:27.940 just have a new parser for it. 71 00:03:29.660 --> 00:03:31.630 Now we're going to build another project as well 72 00:03:31.630 --> 00:03:34.400 in this section, looking at scraping books 73 00:03:34.400 --> 00:03:38.010 and that's gonna be a larger project with a bit more sort 74 00:03:38.010 --> 00:03:41.440 of logging and changing pages and things like that 75 00:03:41.440 --> 00:03:43.080 so it's gonna be a bit longer, 76 00:03:43.080 --> 00:03:46.350 so don't worry too much if not everything is crystal clear, 77 00:03:46.350 --> 00:03:47.960 we're gonna do this again 78 00:03:47.960 --> 00:03:51.770 and we're gonna do more in the books project. 79 00:03:51.770 --> 00:03:53.720 So I'll see you in the very next video.