1

00:00:00,190  -->  00:00:06,960
Welcome to 11:39 which roughly covers pages 237 to 240 of the automate the boring stuff with Python

2

00:00:06,960  -->  00:00:11,660
textbook in this lesson we're going to learn how to download files from the web.

3

00:00:11,880  -->  00:00:16,560
The request module lets you easily download files from the web without having to worry about complicated

4

00:00:16,560  -->  00:00:24,180
issues such as network errors connection problems and data compression and the request module doesn't

5

00:00:24,180  -->  00:00:25,080
come with Python.

6

00:00:25,080  -->  00:00:29,000
It's a third party module so you'll have to install it on your own first.

7

00:00:29,220  -->  00:00:35,930
The course notes and Appendix A have additional details on how to install third party modules.

8

00:00:36,660  -->  00:00:41,490
Now after installing the module do a simple test to make sure the request module is installed correctly

9

00:00:41,490  -->  00:00:41,640
.

10

00:00:41,640  -->  00:00:47,370
Try to import the request module and if no error messages show up then the request module has been successfully

11

00:00:47,370  -->  00:00:57,300
installed to download a file called requests that get ampacity a string of the URL of the file to download

12

00:00:57,300  -->  00:00:58,470
.

13

00:00:58,470  -->  00:01:02,980
Here it automate the boring stuff dot com slash file slash RJ touchbacks T.

14

00:01:03,010  -->  00:01:08,660
I have a text file of the complete text of the Romeo and Juliet play by William Shakespeare.

15

00:01:08,680  -->  00:01:10,230
Let's go ahead just copy this.

16

00:01:10,230  -->  00:01:16,590
By pressing control C and then pressing control V to paste it into idle and the get function returns

17

00:01:16,590  -->  00:01:20,560
a response object which will store in a variable named rez.

18

00:01:20,580  -->  00:01:25,370
This response object contains the response at the web server gave you for this request.

19

00:01:25,590  -->  00:01:30,600
You can tell that the request for the web page succeeded by checking the status code

20

00:01:33,720  -->  00:01:36,640
attribute of the response object.

21

00:01:36,990  -->  00:01:41,880
You may already be familiar with the 404 status code for a file not found.

22

00:01:41,880  -->  00:01:45,980
The 200 status code means everything went OK.

23

00:01:46,110  -->  00:01:52,080
So if the request succeeded the downloaded web pages stored as a string in the response object's text

24

00:01:52,080  -->  00:01:56,280
variable it was variable holds a large string of the entire play.

25

00:01:56,280  -->  00:02:03,720
We passed this to blend in C that this string has a 170 four thousand characters in it.

26

00:02:03,720  -->  00:02:10,710
So let's just print out say the first 500 characters using a slice.

27

00:02:10,950  -->  00:02:16,200
No simpler way to check for success is to call the raise for status method

28

00:02:19,710  -->  00:02:21,360
on the response object.

29

00:02:21,420  -->  00:02:26,460
This will raise an exception if there was an error downloading the file and do nothing if the download

30

00:02:26,460  -->  00:02:27,710
succeeded.

31

00:02:27,810  -->  00:02:28,810
So I call it here.

32

00:02:28,830  -->  00:02:30,940
Nothing happens because it was just fine.

33

00:02:31,280  -->  00:02:35,430
But let's say I had a request I'd get call

34

00:02:40,200  -->  00:02:41,840
to a file that didn't exist.

35

00:02:41,850  -->  00:02:44,470
I'm just going to type some random characters here.

36

00:02:44,490  -->  00:02:49,550
This is a completely random nonexisting file.

37

00:02:49,650  -->  00:02:56,670
I can then call raise for status on this response object and it will raise an exception.

38

00:02:56,670  -->  00:02:59,790
You can see it gives you details about what went wrong right here.

39

00:02:59,820  -->  00:03:05,370
It's giving you a set 404 not found error so the for status method is a good way to ensure that the

40

00:03:05,370  -->  00:03:08,880
program halts if the if a bad download occurs.

41

00:03:08,880  -->  00:03:09,970
This is a good thing.

42

00:03:10,110  -->  00:03:13,950
You want your program to stop as soon as some unexpected error happens.

43

00:03:13,950  -->  00:03:19,950
If a failed download isn't a deal breaker for your program you can wrap the raise for status line inside

44

00:03:19,950  -->  00:03:20,650
of a try.

45

00:03:20,670  -->  00:03:24,400
Except statement to handle this error case without crashing.

46

00:03:24,660  -->  00:03:29,700
From here you can save the web page to a file on your hard drive with the standard open function.

47

00:03:29,700  -->  00:03:31,230
There are some slight differences though.

48

00:03:31,230  -->  00:03:36,900
First you must open the file in right binary mode bypassing the string WB as the second argument to

49

00:03:36,900  -->  00:03:38,560
open.

50

00:03:39,170  -->  00:03:46,740
So I'll just create a file object here and store it in play file call open save it to a file called

51

00:03:46,740  -->  00:03:53,060
Romeo and Juliet up the XTi and open it and write binary mode by passing WB.

52

00:03:53,310  -->  00:03:58,500
Even if the page is in plain text such as the Romeo and Juliet text you downloaded earlier you need

53

00:03:58,500  -->  00:04:04,530
to write binary data instead of text data in order to maintain the Unicode encoding of the text.

54

00:04:04,530  -->  00:04:10,440
Unicode is beyond the scope of this course but there is an excellent explanation of python and Unicode

55

00:04:10,750  -->  00:04:15,840
at bitterly slash unit pain the steps I'm sure you here work for.

56

00:04:15,840  -->  00:04:21,360
Any file that you download from the web using requests to write the web page to a file you can use a

57

00:04:21,360  -->  00:04:26,660
for loop with the response objects itor content method.

58

00:04:26,680  -->  00:04:36,300
I'll just say for chunk in rez content and the content method returns chunks of the content on each

59

00:04:36,300  -->  00:04:41,520
iteration through the loop each chunk of the bytes data type and you get to specify how many bytes each

60

00:04:41,520  -->  00:04:43,160
chunk will contain.

61

00:04:43,620  -->  00:04:50,930
Let's just pass a 100000 100000 bytes per iteration through this loop is pretty good.

62

00:04:52,470  -->  00:04:57,650
So here we just call play file up right and then write out that chunk.

63

00:04:57,770  -->  00:05:02,460
So each chunk is just a piece of that downloaded file in the response object

64

00:05:05,580  -->  00:05:10,050
and the right method will return an integer of how many bytes it wrote to this file.

65

00:05:10,230  -->  00:05:15,480
So a hundred thousand on the first iteration through this loop and on the second iteration it wrote

66

00:05:15,510  -->  00:05:19,820
the remaining seventy four thousand bytes to the file.

67

00:05:19,980  -->  00:05:23,520
We can just call play file that close the file.

68

00:05:23,520  -->  00:05:30,030
Romeo and Juliet tactics t will now exist in the current working directory on your computer.

69

00:05:30,030  -->  00:05:37,590
Note that the file name on the Web site was Arjay dirty XTi but the file on your hard drive has a different

70

00:05:37,590  -->  00:05:40,050
name because we just specified a different name right here.

71

00:05:40,050  -->  00:05:47,310
The open function request module simply handles and downloading the contents of web pages once the pages

72

00:05:47,310  -->  00:05:52,470
downloaded it is simply data in your program like any other variable and that's all there is to the

73

00:05:52,470  -->  00:05:55,050
request module for doing simple downloading.

74

00:05:55,050  -->  00:06:00,590
You can learn about the other features in the request modules from the Web site request dud.

75

00:06:00,690  -->  00:06:02,500
Read the docs dot org.

76

00:06:03,080  -->  00:06:04,080
And one last bit.

77

00:06:04,080  -->  00:06:09,690
The request module is good for downloading files or web pages when you have the exact you are l to download

78

00:06:09,710  -->  00:06:09,850
.

79

00:06:10,020  -->  00:06:15,390
But if you have to log into a website first or figuring out you are ill it is a complicated process

80

00:06:15,390  -->  00:06:15,600
.

81

00:06:15,750  -->  00:06:18,990
Using a request might not be the best way to do it in the next lesson.

82

00:06:19,050  -->  00:06:25,450
We'll learn about selenium which lets your Python scripts control the web browser directly.

83

00:06:25,650  -->  00:06:32,670
To recap the request module is a third party module for downloading web pages and files requests that

84

00:06:32,700  -->  00:06:34,750
get returns a response object.

85

00:06:34,920  -->  00:06:39,720
The raise for status response method will raise an exception if the download failed for some reason

86

00:06:40,350  -->  00:06:45,610
and you can save a downloaded file to your hard drive with calls to the Eidur content method.
