1
00:00:05,200 --> 00:00:10,080
In Python 2, you had to explicitly specify that 
a string is encoded using one of the encodings 

2
00:00:10,080 --> 00:00:13,920
specified in the Unicode standard. 
You had to prefix the string with 

3
00:00:13,920 --> 00:00:19,528
the letter "u", to create a Unicode string.
In Python 3 – which is what we're using –

4
00:00:19,528 --> 00:00:24,800
all strings are Unicode. That makes life a lot easier, 
because we don't have to concern ourselves with 

5
00:00:24,800 --> 00:00:29,760
converting between ASCII and our Python strings.
If you attempt to read a file containing text 

6
00:00:29,760 --> 00:00:33,836
using a different script, or 
alphabet, Python supports that. 

7
00:00:33,920 --> 00:00:37,174
But you do have to be aware 
of the encoding that was used. 

8
00:00:37,360 --> 00:00:40,862
I'll demonstrate the problem, 
then we'll see how to fix it. 

9
00:00:42,240 --> 00:00:46,593
Before I continue, note that you may not 
get the problem on your operating system. 

10
00:00:46,800 --> 00:00:50,538
If you're using Linux or a Mac, 
you probably won't see the problem. 

11
00:00:50,640 --> 00:00:53,840
But I can't say for sure, because 
we haven't configured an operating 

12
00:00:53,840 --> 00:00:56,720
system for every possible 
language that's available. 

13
00:00:56,720 --> 00:00:59,655
You also might not get the 
problem on Windows. 

14
00:00:59,655 --> 00:01:06,288
Again, it depends on the locale (the language, etc.) 
that your operating system is configured for. 

15
00:01:06,400 --> 00:01:11,040
The point is, just because something seems to work 
on your computer, that doesn't mean it will work 

16
00:01:11,040 --> 00:01:15,252
on your users' computers.
Ok, back to the code. 

17
00:01:16,640 --> 00:01:22,103
Open the read_poem.py file, and delete 
everything except the last block of code: 

18
00:01:24,281 --> 00:01:28,720
We've got a simple loop, that prints 
each line in the Jabberwocky text file. 

19
00:01:28,720 --> 00:01:31,932
We made the loop break after printing 
the line containing "jubjub", 

20
00:01:31,932 --> 00:01:39,840
so I'll delete those last 2 lines as well.
Ok, let's run the program, and see the problem. 

21
00:01:40,800 --> 00:01:43,532
The last line of the file contains 
an m-dash character, 

22
00:01:43,532 --> 00:01:48,240
followed by the name of the author – Lewis Carroll.
That's a common way to include the source of a 

23
00:01:48,240 --> 00:01:54,889
quote in text. But as you can see, in my output, I 
get three strange characters instead of the dash. 

24
00:01:54,960 --> 00:02:01,144
Scrolling up, we get the same problem in the third
stanza, after "Long time the manxome foe he sought". 

25
00:02:01,144 --> 00:02:02,640
So what's going on? 

26
00:02:02,640 --> 00:02:06,960
I've already said that the files in the resources 
are encoded as UTF-8, and if you've read the 

27
00:02:06,960 --> 00:02:12,880
documentation you'll have seen comments like
"The default encoding used by Python is UTF-8", 

28
00:02:12,880 --> 00:02:16,400
so why isn't this working?
The reason is, that comment in 

29
00:02:16,400 --> 00:02:21,688
the documentation is often mis-interpreted. It 
means that Python expects source files – 

30
00:02:21,688 --> 00:02:25,920
your code files – to be in UTF-8.
It doesn't mean that Python will 

31
00:02:25,920 --> 00:02:30,596
automatically try to read your text 
files as UTF-8, when you open them. 

32
00:02:30,640 --> 00:02:35,237
I'll fix the problem, then we'll have another 
look at the documentation for the open function. 

33
00:02:35,360 --> 00:02:39,147
To fix the problem, we specify 
the encoding explicitly: 

34
00:02:44,800 --> 00:02:50,466
Ok, we've told the open function that our 
Jabberwocky.txt file is encoded using UTF-8. 

35
00:02:50,560 --> 00:02:55,537
Run the program:
and we get the correct output. 

36
00:02:55,600 --> 00:02:58,626
The contents of the file are 
now being interpreted correctly. 

37
00:02:58,800 --> 00:03:04,620
Let's have another look at the documentation for 
the open function. I'll open it in my browser. 

38
00:03:08,880 --> 00:03:11,680
The relevant text is in the 3rd paragraph: 

39
00:03:11,680 --> 00:03:16,542
In text mode, if encoding is not 
specified the encoding used is platform dependent: 

40
00:03:16,542 --> 00:03:22,031
locale.getpreferredencoding(False) is 
called to get the current locale encoding 

41
00:03:23,360 --> 00:03:28,669
If the encoding is "platform dependent", you will 
get different behaviour on different computers. 

42
00:03:28,720 --> 00:03:32,720
That's not good. We don't want to write 
software that will work for some people, 

43
00:03:32,720 --> 00:03:36,080
and fail for others!
When opening a text file, 

44
00:03:36,080 --> 00:03:38,454
explicitly specify the encoding. 

45
00:03:38,640 --> 00:03:41,912
That applies whether you're opening 
the file for reading or writing. 

46
00:03:44,320 --> 00:03:48,560
When you're writing to a file, you can 
choose which encoding you want to use. 

47
00:03:48,560 --> 00:03:52,400
Normally, you'd use UTF-8, if it 
supports the character set – 

48
00:03:52,400 --> 00:03:55,680
the alphabet – that you want to use.
But how do you know what 

49
00:03:55,680 --> 00:04:06,160
encoding to specify, when reading a file?
We'll see some ways to do that, in the next video.

