1
00:00:05,120 --> 00:00:09,920
When you open a file, in Python, you 
tell it what mode to open the file in. 

2
00:00:09,920 --> 00:00:16,320
Up to now, we've seen the "r" mode, to read from 
the file, and the "w" mode, to write to the file. 

3
00:00:16,320 --> 00:00:21,280
We've also implicitly used the default text 
mode. That's the default for the type of file, 

4
00:00:21,280 --> 00:00:25,600
so we've been working with text files.
Let's have a look at the documentation, 

5
00:00:25,600 --> 00:00:29,520
to see the various modes available.
The documentation will also answer 

6
00:00:29,520 --> 00:00:34,930
the question we had, at the end of the previous 
video. "Where did our first set of numbers go?". 

7
00:00:34,930 --> 00:00:39,636
I'll start by opening the documentation 
for the open function, in my browser. 

8
00:00:40,080 --> 00:00:44,880
So far, we've used the first two arguments; 
we provided a file name for the file argument, 

9
00:00:44,880 --> 00:00:51,280
and either "r" or "w" for the mode.
The table summarises the various modes available. 

10
00:00:51,280 --> 00:00:55,440
"r" opens a file for reading, and 
is the default. We didn't include 

11
00:00:55,440 --> 00:01:00,080
"r" a couple of times, and the file still 
opened for reading without any problems. 

12
00:01:00,080 --> 00:01:05,600
"w" opens a file for writing – as we've seen.
When you're using documentation like this, 

13
00:01:05,600 --> 00:01:09,360
read it carefully. Although the 
table does mention that "w" will 

14
00:01:09,360 --> 00:01:14,960
truncate the file first, that is only a summary.
You'll want to read the paragraph above the table, 

15
00:01:14,960 --> 00:01:18,000
to understand the complete 
implications of that note. 

16
00:01:18,080 --> 00:01:23,120
Other common values are 'w' for writing 
(truncating the file if it already exists) ... 

17
00:01:23,120 --> 00:01:27,920
That's why the previous numbers weren't in our 
file. Each time we opened the file for writing, 

18
00:01:27,920 --> 00:01:32,769
the existing contents were deleted. 
We start with an empty file each time. 

19
00:01:32,769 --> 00:01:39,120
So mode "w" will create a file, if it doesn't 
already exist, and clear it out, if it does exist. 

20
00:01:39,120 --> 00:01:44,381
In either case – whether the file already existed 
or not – you start off with an empty file. 

21
00:01:44,560 --> 00:01:48,240
If you expect that the file doesn't exist, 
but you want to make sure you don't clear it 

22
00:01:48,240 --> 00:01:53,200
out if it does, you can use mode "x". That 
will create a new file, but will raise an 

23
00:01:53,200 --> 00:01:57,776
error if the file already exists.
You probably won't use mode "x" often, 

24
00:01:57,776 --> 00:02:01,200
but it's there if you need it.
Mode "x" will guard against accidentally 

25
00:02:01,200 --> 00:02:06,160
deleting the contents of an existing file. But 
you'll need to catch the error, if you don't want 

26
00:02:06,160 --> 00:02:11,200
your code to crash, when the file does exist.
If you want to add more text to the end of an 

27
00:02:11,200 --> 00:02:16,160
existing file, you'd specify "a" for the 
mode. That will create a new file if one 

28
00:02:16,160 --> 00:02:20,908
doesn't already exist, and appends to the 
end of the existing file, if there is one. 

29
00:02:21,040 --> 00:02:25,157
It's similar to mode "w", but 
doesn't truncate the file. 

30
00:02:25,280 --> 00:02:29,834
When dealing with text files, those are the 
only modes you'll usually be interested in. 

31
00:02:30,000 --> 00:02:33,720
It's very rare that you'd want to open a 
text file for both reading and writing. 

32
00:02:33,840 --> 00:02:37,468
One example might be if you've numbered 
each line of the file in some way. 

33
00:02:37,600 --> 00:02:43,147
In that case, you'd need to read the number of the 
last line, before writing more lines to the file. 

34
00:02:43,200 --> 00:02:47,280
If you did need to do that, you 
could specify a mode of "r+". 

35
00:02:47,440 --> 00:02:52,640
If you use "w+" for the mode, the file will 
be truncated when you open it. You can read 

36
00:02:52,640 --> 00:02:57,440
anything that you write after opening the file, 
but initially it will be empty. It's very hard 

37
00:02:57,440 --> 00:03:03,360
to think of a use for that mode with text files.
Ok, as you can see, there is a lot more you could 

38
00:03:03,360 --> 00:03:07,840
read up about, when dealing with files.
Most of it is quite advanced, 

39
00:03:07,840 --> 00:03:11,605
and unlikely to be useful unless you're 
doing some pretty low-level stuff. 

40
00:03:11,760 --> 00:03:16,804
We've covered almost everything you need to know, 
to read and write text files in your programs. 

41
00:03:16,804 --> 00:03:19,360
There are a couple more things I 
want to talk about, with the read 

42
00:03:19,360 --> 00:03:23,120
and readlines methods that we've used.
We'll also be looking at two more of the 

43
00:03:23,120 --> 00:03:28,080
parameters you can use with the open function. If 
you've scrolled this page, scroll back up until 

44
00:03:28,080 --> 00:03:33,431
you can see the function signature for open.
The third parameter is buffering. 

45
00:03:34,640 --> 00:03:37,955
Because disk drives are slow 
– compared to computer memory, 

46
00:03:37,955 --> 00:03:42,852
they're very slow – data is read in large 
chunks into an area of memory called a "buffer". 

47
00:03:42,960 --> 00:03:46,400
That makes the data available when you 
want to read more data, without having to 

48
00:03:46,400 --> 00:03:51,920
wait for the disk drive to spin around again.
Of course, if you're using modern SSD drives, 

49
00:03:51,920 --> 00:03:56,560
they don't spin. But they're still a 
lot slower than your computer's memory. 

50
00:03:56,560 --> 00:03:59,600
By reading larger chunks from the 
drive, the next data that you'll 

51
00:03:59,600 --> 00:04:05,042
probably want is made available faster.
A disk buffer is also called a disk cache. 

52
00:04:07,440 --> 00:04:11,547
When writing data to disk, the 
data is put into a buffer first. 

53
00:04:11,600 --> 00:04:16,404
That allows your program to move on, without 
waiting for the actual transfer to complete. 

54
00:04:16,480 --> 00:04:20,952
The write to disk can happen in the background, 
while your code gets on with other things. 

55
00:04:21,040 --> 00:04:25,600
That's another reason why it's very important 
to always close a file, when writing. 

56
00:04:25,600 --> 00:04:28,960
If you fail to close the file, 
and your program terminates, 

57
00:04:28,960 --> 00:04:33,624
data in the cache might not be written 
to the disk – resulting in data loss. 

58
00:04:34,880 --> 00:04:39,468
I won't be talking about the remaining 
parameters, except for encoding and newline. 

59
00:04:39,600 --> 00:04:43,593
We'll provide a value for newline 
when we work with CSV files, shortly, 

60
00:04:43,593 --> 00:04:48,212
so I'll leave that one until then.
It is important that you understand what encoding is, 

61
00:04:48,212 --> 00:04:52,560
because it can be a common source of problems.
In order for encoding to make sense, 

62
00:04:52,560 --> 00:04:56,320
we have to dig a bit deeper into 
Unicode and what encoding means. 

63
00:04:56,320 --> 00:04:58,720
I'll see you in the next video.

