WEBVTT 1 00:00:00.790 --> 00:00:02.640 Hi guys and welcome back. 2 00:00:02.640 --> 00:00:05.150 In this section, we're gonna have a Stata code, 3 00:00:05.150 --> 00:00:07.360 which is the code from the last section, 4 00:00:07.360 --> 00:00:08.930 the web scraping code. 5 00:00:08.930 --> 00:00:10.360 Let's go over it real quick, 6 00:00:10.360 --> 00:00:13.050 just in case you haven't seen it in a while. 7 00:00:13.050 --> 00:00:16.310 We've got the entry point to our code, app.py, 8 00:00:16.310 --> 00:00:20.030 and that makes the requests to quotes.toscrape.com, 9 00:00:20.030 --> 00:00:21.910 which is just a quotes website 10 00:00:21.910 --> 00:00:23.690 that contains some quotes by famous people 11 00:00:23.690 --> 00:00:26.300 and gives us the HTML content of the page. 12 00:00:26.300 --> 00:00:30.063 That HTML content is passed on to our QuotesPage class. 13 00:00:30.063 --> 00:00:33.360 What that does is it creates a BeautifulSoup object, 14 00:00:33.360 --> 00:00:36.780 which will allow us to traverse the HTML that we've received 15 00:00:36.780 --> 00:00:39.470 and find things inside the HTML, 16 00:00:39.470 --> 00:00:42.630 which is what we do down here in the Quote property. 17 00:00:42.630 --> 00:00:45.470 We go into that BeautifulSoup object. 18 00:00:45.470 --> 00:00:48.850 And we find elements, using the select method, 19 00:00:48.850 --> 00:00:51.370 matching this selector here. 20 00:00:51.370 --> 00:00:55.210 The selector is just any div in the page 21 00:00:55.210 --> 00:00:57.890 that has the class, quote. 22 00:00:57.890 --> 00:01:00.020 If you haven't seen the last section, 23 00:01:00.020 --> 00:01:02.160 I would recommend going through that first, 24 00:01:02.160 --> 00:01:04.870 just because it contains a bunch of explanations 25 00:01:04.870 --> 00:01:07.540 that you will need for this section as well. 26 00:01:07.540 --> 00:01:10.620 Now, after finding all the divs with that class with quote, 27 00:01:10.620 --> 00:01:15.280 it passes them into a QuoteParser object that we also made. 28 00:01:15.280 --> 00:01:17.580 And this one stores the parent, 29 00:01:17.580 --> 00:01:20.550 which is basically the container of the quote. 30 00:01:20.550 --> 00:01:23.380 And then, has if you properties, as well; 31 00:01:23.380 --> 00:01:25.710 content, author, and tags. 32 00:01:25.710 --> 00:01:27.930 The content property, once again, 33 00:01:27.930 --> 00:01:31.210 gets a locator from QuoteLocators file in this case 34 00:01:31.210 --> 00:01:34.410 and then, finds an element inside 35 00:01:34.410 --> 00:01:38.090 this quote with that locator. 36 00:01:38.090 --> 00:01:39.730 Same for author and tags. 37 00:01:39.730 --> 00:01:41.830 So basically, what is gonna happen is 38 00:01:41.830 --> 00:01:44.900 here we're gonna get a bunch of QuoteParcers, 39 00:01:44.900 --> 00:01:47.050 one for each div, 40 00:01:47.050 --> 00:01:48.670 and then, each of these objects, 41 00:01:48.670 --> 00:01:50.630 we can use to get the content author 42 00:01:50.630 --> 00:01:52.303 and tags of that quote. 43 00:01:54.050 --> 00:01:56.900 Then, at the end, we print out the page quotes. 44 00:01:56.900 --> 00:01:59.160 So you can actually run this just by right-clicking 45 00:01:59.160 --> 00:02:01.580 and running app.py and it's gonna launch 46 00:02:01.580 --> 00:02:05.120 and it's gonna give us all these quotes that we've got. 47 00:02:05.120 --> 00:02:07.320 So that has actually launched the page. 48 00:02:07.320 --> 00:02:09.630 Well, not launched the page, but it's gone to the page 49 00:02:09.630 --> 00:02:13.290 and got the HTML content using the requests module. 50 00:02:13.290 --> 00:02:15.350 And once we can the HTML content, 51 00:02:15.350 --> 00:02:18.073 it returned it to us and we've performed our searches. 52 00:02:18.910 --> 00:02:19.940 There's a problem, of course, 53 00:02:19.940 --> 00:02:24.940 which is that if we use JavaScript, then this won't work. 54 00:02:25.740 --> 00:02:27.820 Or if we were to navigate any other page 55 00:02:27.820 --> 00:02:29.690 that uses JavaScript, this won't work. 56 00:02:29.690 --> 00:02:31.460 I'll show you an example. 57 00:02:31.460 --> 00:02:34.610 Here, I've launched another quotes.toscrape.com instance, 58 00:02:34.610 --> 00:02:38.440 but this one is /search.aspx. 59 00:02:38.440 --> 00:02:40.890 This is a slightly different page that uses JavaScript, 60 00:02:40.890 --> 00:02:45.200 so we can use this to try out our JavaScript scrapping. 61 00:02:45.200 --> 00:02:48.010 Notice that, if we were to use the requests module 62 00:02:48.010 --> 00:02:50.850 to get the HTML content of this page, 63 00:02:50.850 --> 00:02:52.970 we would get what the inspector 64 00:02:52.970 --> 00:02:54.630 gives us at this point in time. 65 00:02:54.630 --> 00:02:58.080 So we've got a couple drop-downs 66 00:02:58.080 --> 00:03:00.190 and a search button that is now partially hidden. 67 00:03:00.190 --> 00:03:01.620 I will move that a bit. 68 00:03:01.620 --> 00:03:03.490 There you go, my search button down there. 69 00:03:03.490 --> 00:03:08.490 Now that has a problem because the author drop-down, 70 00:03:10.780 --> 00:03:14.110 that you can see, here has a bunch of authors inside it. 71 00:03:14.110 --> 00:03:15.550 You can see them if you click on it 72 00:03:15.550 --> 00:03:17.750 and you can see it in the HTML. 73 00:03:17.750 --> 00:03:21.570 But the tag drop-down does not have any elements. 74 00:03:21.570 --> 00:03:26.110 It is only when you select an author that the tags change 75 00:03:26.110 --> 00:03:28.690 for tags that are available for this author. 76 00:03:28.690 --> 00:03:31.850 So basically, you're gonna be able to only select tags 77 00:03:31.850 --> 00:03:35.510 that match that author or things that the author sent. 78 00:03:35.510 --> 00:03:38.523 Because that change happens after the page is loaded, 79 00:03:39.690 --> 00:03:41.990 it means that our requests code 80 00:03:41.990 --> 00:03:45.210 cannot deal with something like this for two reasons. 81 00:03:45.210 --> 00:03:48.880 The first one is that the content of the page 82 00:03:48.880 --> 00:03:50.860 is changing using JavaScript. 83 00:03:50.860 --> 00:03:53.080 That is the only way to change page content 84 00:03:53.080 --> 00:03:55.320 and when we load a page using requests, 85 00:03:55.320 --> 00:03:58.730 we cannot execute JavaScript, so that is a problem. 86 00:03:58.730 --> 00:04:01.750 The second problem is we cannot use the requests module 87 00:04:01.750 --> 00:04:05.790 to interact with a drop-down inside the page. 88 00:04:05.790 --> 00:04:08.670 The requests module allows us to get the HTML content, 89 00:04:08.670 --> 00:04:11.210 so we would be able to see there is a drop-down, 90 00:04:11.210 --> 00:04:14.750 but we won't be able to click on it and make a change, 91 00:04:14.750 --> 00:04:17.060 so the requests code is not going 92 00:04:17.060 --> 00:04:20.190 to allow us to use this page. 93 00:04:20.190 --> 00:04:21.440 Then, if you click the search button, 94 00:04:21.440 --> 00:04:23.690 a further JavaScript change happens. 95 00:04:23.690 --> 00:04:27.830 That shows you a quote matching the results. 96 00:04:27.830 --> 00:04:29.610 What we have to do is we have to actually 97 00:04:29.610 --> 00:04:32.900 launch the Chrome browser, navigate to this page, 98 00:04:32.900 --> 00:04:35.500 and modify these drop-downs ourselves, 99 00:04:35.500 --> 00:04:37.160 then click the search button 100 00:04:37.160 --> 00:04:39.493 and finally, extract the result. 101 00:04:41.070 --> 00:04:43.380 So that's what we're gonna be doing in this section. 102 00:04:43.380 --> 00:04:47.030 As I mentioned earlier on, we've seen the code as it is now 103 00:04:47.030 --> 00:04:49.870 and that code as it is now is just for this page here, 104 00:04:49.870 --> 00:04:51.510 which doesn't use JavaScript. 105 00:04:51.510 --> 00:04:52.790 For our new page, 106 00:04:52.790 --> 00:04:55.730 we're gonna be doing everything in this section, 107 00:04:55.730 --> 00:04:57.353 so I'll see you in the next video.