1
00:00:05,280 --> 00:00:08,880
Ok, let's have a look at the 
encodings for our text files. 

2
00:00:08,880 --> 00:00:14,400
Open Jabberwocky.txt in IntelliJ, and have a 
look at the right hand side of the status bar. 

3
00:00:14,400 --> 00:00:18,720
There's some useful information down there. 
As you move the cursor around the document, 

4
00:00:18,720 --> 00:00:23,520
you can see which line and column you're on.
If you select text, that changes to show 

5
00:00:23,520 --> 00:00:27,360
how many characters are selected – which 
can save you having to count characters, 

6
00:00:27,360 --> 00:00:32,159
when you want to know how long a string is.
If you don't see all of those indicators, 

7
00:00:32,159 --> 00:00:36,160
go to the View menu, and choose 
Appearance > Status Bar Widgets, 

8
00:00:36,160 --> 00:00:40,634
and tick File Encoding (or whichever 
other indicator you want to display). 

9
00:00:40,720 --> 00:00:44,560
To the right of the line and 
column numbers, I've got CRLF. 

10
00:00:44,560 --> 00:00:49,237
That indicates the characters that are 
used to represent a newline. On Windows, 

11
00:00:49,237 --> 00:00:54,240
a newline is represented by two characters: 
a carriage return followed by a line feed. 

12
00:00:54,240 --> 00:00:59,840
On Mac and Linux computers, a single line feed 
character is used instead. If you're using Linux 

13
00:00:59,840 --> 00:01:05,600
or a Mac, you should see LF there, in place 
of the CRLF that my Windows computer is using. 

14
00:01:05,600 --> 00:01:10,560
We don't normally have to worry about the 
newline code, because Python handles that for us. 

15
00:01:10,560 --> 00:01:13,920
Python will send a carriage return 
and line feed on Windows computers, 

16
00:01:13,920 --> 00:01:17,199
and a line feed by itself on Linux or Mac. 

17
00:01:17,280 --> 00:01:21,520
Moving on, the next indicator on the 
status line is the file encoding. 

18
00:01:21,520 --> 00:01:26,628
When you open a text file in IntelliJ, it 
does its best to work out the file encoding. 

19
00:01:26,720 --> 00:01:29,360
Which answers the question from 
the last video: "How do you know 

20
00:01:29,360 --> 00:01:35,840
what encoding to specify, when reading a file?".
The simple answer is, open the file in IntelliJ. 

21
00:01:35,840 --> 00:01:39,760
You can also use Notepad, the 
editor that comes with Windows. 

22
00:01:39,760 --> 00:01:44,080
On Windows 10, that has a status 
bar similar to IntelliJ's. 

23
00:01:44,080 --> 00:01:48,000
So our Jabberwocky.txt file 
is encoded using UTF-8. 

24
00:01:48,000 --> 00:01:52,000
That was easy – it didn't need an 
entire video just to show you that! 

25
00:01:52,000 --> 00:01:56,800
I've obviously got something else planned.
These indicators are clickable, and let you 

26
00:01:56,800 --> 00:02:02,160
change the newline characters, and the encoding.
You can also jump directly to a particular line 

27
00:02:02,160 --> 00:02:06,960
and column in the file, by clicking the 
line:column indicator. I won't do that, 

28
00:02:06,960 --> 00:02:11,200
click it yourself and experiment, if that's 
something you think you'll find useful. 

29
00:02:11,200 --> 00:02:15,947
I will change the encoding though. That will let 
us experiment with different encodings, 

30
00:02:15,947 --> 00:02:21,200
and see the sort of error you can get, if you specify an 
incompatible encoding for the file you're reading. 

31
00:02:21,200 --> 00:02:23,840
Click UTF-8 on the status 
bar, and you'll get a list 

32
00:02:23,840 --> 00:02:30,468
of all the encodings that IntelliJ recognises.
At first, it only shows the most common encodings: 

33
00:02:30,560 --> 00:02:34,560
ISO-8859-1
US-ASCII 

34
00:02:34,560 --> 00:02:40,126
UTF-16
and windows-1252 (on Windows only) 

35
00:02:41,200 --> 00:02:47,520
ISO-8859-1 was the basis for the first 
256 Unicode code points. These are all 

36
00:02:47,520 --> 00:02:53,040
single-byte representations of characters, with 
the first 127 matching the ASCII characters. 

37
00:02:53,040 --> 00:02:59,040
If you google ISO-8859-1, you'll find a table on 
the Wikipedia page, showing the characters that 

38
00:02:59,040 --> 00:03:05,720
are represented. It includes most of the accented 
characters in use in Western European languages. 

39
00:03:07,600 --> 00:03:11,962
Code point refers to the individual codes 
that are defined to represent characters. 

40
00:03:11,962 --> 00:03:16,800
Although the term wasn't used back then, I'll 
use ASCII to explain what a code point is. 

41
00:03:16,800 --> 00:03:21,680
As we've seen, ASCII can represent 
127 different "characters". 

42
00:03:21,680 --> 00:03:27,840
Note that we don't count zero in that range.
The code point 65 represents a capital letter A. 

43
00:03:27,840 --> 00:03:32,080
The first 32 code points are control 
codes – things like 9 for a tab, 

44
00:03:32,080 --> 00:03:35,840
10 for a line feed, and 13 for a carriage return. 

45
00:03:35,840 --> 00:03:39,685
Back then, they were called ASCII 
codes rather than code points. 

46
00:03:39,760 --> 00:03:45,085
But when dealing with more modern encodings, 
we talk about code points instead. 

47
00:03:46,720 --> 00:03:51,520
When dealing with code points, note that 
they're not restricted to single-byte values. 

48
00:03:51,520 --> 00:03:56,400
If they were, we'd still be stuck 
with only 256 possible codes. 

49
00:03:56,400 --> 00:04:02,880
Instead, a code point can be 1 or more bytes.
UTF-8, for example, uses 1 byte for the first 

50
00:04:02,880 --> 00:04:08,393
128 characters, which corresponds 
to the ASCII character set. 

51
00:04:09,760 --> 00:04:14,240
The next 1,920 characters 
are represented by 2 bytes. 

52
00:04:14,240 --> 00:04:20,560
Following on, and using 3 bytes each, are most 
of the Chinese, Japanese and Korean characters. 

53
00:04:20,560 --> 00:04:26,160
Lesser used CJK (Chinese, Japanese, 
Korean) characters, mathematical symbols, 

54
00:04:26,160 --> 00:04:31,455
musical characters and emojis 
are encoded using 4 bytes each. 

55
00:04:33,600 --> 00:04:36,559
US-ASCII is another name for ASCII. 

56
00:04:38,720 --> 00:04:44,245
UTF-16 is a Unicode Transformation 
Format that uses one or two 16 bit codes. 

57
00:04:44,245 --> 00:04:49,760
It was adopted by Microsoft, but 
since 2019 Windows switched to UTF-8. 

58
00:04:49,760 --> 00:04:53,600
If the smallest size of a code 
point is 2 bytes (16 bits), 

59
00:04:53,600 --> 00:04:59,008
a file using only the Latin alphabet will be twice 
as large as the same text encoded with UTF-8. 

60
00:04:59,008 --> 00:05:04,202
That's one reason why UTF-16 
has fallen out of favour. 

61
00:05:05,200 --> 00:05:09,360
Windows-1252 is a single-byte 
encoding adopted by Microsoft. 

62
00:05:09,360 --> 00:05:13,600
It's basically ISO-8859-1 
with some extra characters. 

63
00:05:13,600 --> 00:05:17,875
It's important, because that's the encoding 
that Python will most likely choose, 

64
00:05:17,875 --> 00:05:22,880
when you call the open function on Windows.
Of course, this is only true if your Windows setup 

65
00:05:22,880 --> 00:05:27,760
is using the Latin alphabet. If you're using a 
different alphabet for your language, the Windows 

66
00:05:27,760 --> 00:05:33,279
file system will use a suitable encoding, and 
that will be the default encoding used by open. 

67
00:05:34,400 --> 00:05:39,778
Back in IntelliJ, you can see more formats 
that IntelliJ recognises, by clicking More. 

68
00:05:39,778 --> 00:05:42,960
As you can see, there are a lot of encodings. 

69
00:05:42,960 --> 00:05:47,200
IBM, Microsoft and Apple all 
created their own encodings, 

70
00:05:47,200 --> 00:05:53,840
to try to solve the limitations of 7-bit ASCII.
Ok, I'll click UTF-8 again, to clear that list, 

71
00:05:53,840 --> 00:05:57,600
and get the shorter list back.
Those warning and error symbols, 

72
00:05:57,600 --> 00:06:03,120
next to each one, show encodings that will either 
change the contents of the file, or lose data. 

73
00:06:03,120 --> 00:06:08,474
Be aware of that, and take care, if you really 
do want to change the encoding of a file. 

74
00:06:08,560 --> 00:06:14,327
I'm going to change the encoding to windows-1252.
I'm not suggesting that this is something you should do, 

75
00:06:14,327 --> 00:06:17,680
 I just want you to see some of the 
problems that can happen, when Python attempts 

76
00:06:17,680 --> 00:06:21,840
to read a file in the wrong format.
We saw one problem – we got strange 

77
00:06:21,840 --> 00:06:28,000
characters in the text. Most of the text was still 
readable, but that might not always be the case. 

78
00:06:28,000 --> 00:06:32,800
You can also get an error, if the encoding 
you specify doesn't match the file encoding. 

79
00:06:32,800 --> 00:06:37,120
That's what I'll demonstrate now.
So I've chosen windows-1252, 

80
00:06:37,120 --> 00:06:41,600
and IntelliJ is going to change the 
encoding of our Jabberwocky.txt file. 

81
00:06:41,600 --> 00:06:46,729
If you don't see windows-1252 in the short 
list, click More and find it in there. 

82
00:06:47,120 --> 00:06:51,235
We get a warning that changing the encoding 
might change the contents of the file. 

83
00:06:51,360 --> 00:06:56,802
Normally, you'd Cancel at that point. But 
I'm going to go ahead, and click Convert. 

84
00:06:57,120 --> 00:07:02,720
The encoding, at the right hand side 
of the status bar, is now windows-1252. 

85
00:07:02,720 --> 00:07:07,360
I'll use the green Run triangle, on 
the toolbar, to run our program again. 

86
00:07:07,360 --> 00:07:12,720
Notice that it's still set to run read_poem. 
IntelliJ retains the program you ran last, 

87
00:07:12,720 --> 00:07:17,459
for situations like this. When you're switching 
between different files in the editor, 

88
00:07:17,459 --> 00:07:21,506
it's very useful that the Run triangle 
will continue to run your main program. 

89
00:07:21,600 --> 00:07:26,000
Ok, run the program.
And it crashes! 

90
00:07:26,000 --> 00:07:31,840
We've got a UnicodeDecodeError at position 334.
When you see UnicodeDecodeError, 

91
00:07:31,840 --> 00:07:36,007
that almost always means that you've 
specified the wrong encoding for the file. 

92
00:07:36,007 --> 00:07:40,160
The position can be helpful, to know 
which character is causing the problem. 

93
00:07:40,160 --> 00:07:46,160
Position the cursor right at the start of the 
file, and start selecting text. The indicator, 

94
00:07:46,160 --> 00:07:51,280
on the status bar, tells you which row and column 
you're positioned on. It also shows you how many 

95
00:07:51,280 --> 00:07:57,280
characters you've selected, so keep selecting 
text until you've got 334 characters selected. 

96
00:07:57,280 --> 00:08:01,760
You probably guessed where you'd end up, just 
before that m-dash character. That's the one 

97
00:08:01,760 --> 00:08:07,040
that was rendered incorrectly, when we attempted 
to read the file without specifying an encoding. 

98
00:08:07,040 --> 00:08:12,160
The windows-1252 code point for that 
character isn't a valid code point in UTF-8. 

99
00:08:12,160 --> 00:08:17,835
And we've specified UTF-8 as the encoding.
Switch back to read_poem.py. 

100
00:08:17,920 --> 00:08:24,735
On line 1, we set the encoding to "utf-8".
I'll change that to "windows-1252": 

101
00:08:25,760 --> 00:08:27,740
and run the program. 

102
00:08:29,600 --> 00:08:33,039
That's fixed the error, and 
the file can be read correctly. 

103
00:08:33,039 --> 00:08:38,308
If you get problems opening a text file, the 
most common cause will be an incorrect encoding. 

104
00:08:38,308 --> 00:08:40,559
And you now know how to deal with that. 

105
00:08:40,559 --> 00:08:45,000
I'll finish this video by showing you the 
encodings that Python currently supports. 

106
00:08:45,000 --> 00:08:50,000
Note that this isn't the same as the list we saw 
in IntelliJ. IntelliJ is an IDE that we use to 

107
00:08:50,000 --> 00:08:55,360
edit and run our Python code, but it's Python 
that actually interprets and executes the code. 

108
00:08:55,360 --> 00:09:00,481
The encodings are in the codecs module, 
and I'll switch to that page in my browser. 

109
00:09:03,920 --> 00:09:09,925
Click Standard Encodings in the left hand menu, 
to see all the encodings that Python supports. 

110
00:09:11,360 --> 00:09:13,520
The first column gives the name of the codec, 

111
00:09:13,520 --> 00:09:17,040
but you can also use any of the 
aliases from the second column. 

112
00:09:17,040 --> 00:09:21,754
Helpfully, the table also shows the 
languages that each codec supports. 

113
00:09:21,754 --> 00:09:26,480
Scrolling down – or use Ctrl-F to search 
for windows-1252 – that's the encoding 

114
00:09:26,480 --> 00:09:33,520
name we used in read_poem. We could also 
have used cp1252, to get the same result. 

115
00:09:33,520 --> 00:09:36,960
Ok, there's one more thing I 
want to draw your attention to. 

116
00:09:36,960 --> 00:09:43,760
Scroll down to the last entry in the table.
Below utf_8 there's a utf_8_sig entry. 

117
00:09:43,760 --> 00:09:48,240
If you get an error when attempting to read 
a file as UTF-8, and you know it is UTF-8, 

118
00:09:48,240 --> 00:09:53,200
the problem might be caused by a BOM – 
Byte Order Mark – at the start of the file. 

119
00:09:53,200 --> 00:09:58,800
Using a BOM is discouraged for UTF-8, but 
some programs and websites still include one. 

120
00:09:58,800 --> 00:10:02,640
If you create the file using Notepad 
on Windows, for example, you may have 

121
00:10:02,640 --> 00:10:08,560
an invisible BOM at the start of the file.
The utf_8_sig encoding will skip the marker, 

122
00:10:08,560 --> 00:10:13,280
and should fix the problem.
And that's enough about Unicode and encodings. 

123
00:10:13,280 --> 00:10:16,720
In the next few videos, we'll look at 
reading and saving data in a couple 

124
00:10:16,720 --> 00:10:21,900
of common formats: JSON and CSV.
I'll see you in the next video.

