1
00:00:05,040 --> 00:00:11,040
We won't be typing any code in this video, so 
grab a pen and paper, and take some notes instead. 

2
00:00:11,040 --> 00:00:14,640
Back in the last section, when we 
wrote our Really bad hashing function, 

3
00:00:14,640 --> 00:00:18,800
we saw that each character is stored in the 
computer as a number. I mentioned ASCII and 

4
00:00:18,800 --> 00:00:24,783
Unicode at the time. In this video, we'll 
take a closer look at what Unicode is. 

5
00:00:25,840 --> 00:00:30,240
A long long time ago, when the world was young, 
characters were encoded using a single byte 

6
00:00:30,240 --> 00:00:32,960
for each character.
A byte is 8 bits, 

7
00:00:32,960 --> 00:00:36,822
and can represent decimal
numbers from 0 to 255. 

8
00:00:36,880 --> 00:00:41,494
That meant a maximum of 256 
characters could be represented. 

9
00:00:42,400 --> 00:00:47,680
There were 2 different encodings in common use:
IBM mainframes used EBCDIC 

10
00:00:47,680 --> 00:00:54,720
(pronounced ebb-sid-ick). They still do today.
As personal computers developed, they used ASCII. 

11
00:00:54,720 --> 00:01:00,830
ASCII is a 7-bit code – so it could 
only represent 127 different characters. 

12
00:01:02,640 --> 00:01:06,480
ASCII became the most popular 
encoding on PCs and smaller computers, 

13
00:01:06,480 --> 00:01:10,960
and you'll still hear programmers referring to 
ASCII when talking about encoding characters. 

14
00:01:10,960 --> 00:01:15,360
Despite only being able to represent 
127 characters and control codes, 

15
00:01:15,360 --> 00:01:19,250
ASCII was suitable for the 
computational needs of the time. 

16
00:01:20,960 --> 00:01:24,480
But times have changed. Even the 
visionaries of the past couldn't 

17
00:01:24,480 --> 00:01:30,720
predict just how much things would change.
Thomas J. Watson, the founder of IBM, said: 

18
00:01:30,720 --> 00:01:34,080
I think there is a world market 
for maybe five computers. 

19
00:01:34,080 --> 00:01:40,000
MS-DOS, the operating system that the first 
versions of Windows ran on, was restricted to 640 

20
00:01:40,000 --> 00:01:47,860
kilobytes of addressable memory. Bill Gates said:
640K of memory should be enough for everyone. 

21
00:01:49,440 --> 00:01:53,040
ASCII was fine for representing the Latin 
alphabet used in the English language, 

22
00:01:53,040 --> 00:01:57,520
but couldn't represent characters with accents. 
That means ASCII couldn't represent all the 

23
00:01:57,520 --> 00:02:02,930
characters in languages that use the Latin 
alphabet, but also include diacritics (accents). 

24
00:02:03,040 --> 00:02:08,160
All modern European languages, except 
English, commonly use diacritics. 

25
00:02:08,160 --> 00:02:11,918
Any language that didn't use the Latin 
alphabet couldn't be represented at all. 

26
00:02:12,080 --> 00:02:13,920
So most of the world couldn't use a computer in 

27
00:02:13,920 --> 00:02:16,181
their native language.

28
00:02:18,160 --> 00:02:21,440
Technical debt is a term used to 
describe problems that we face, 

29
00:02:21,440 --> 00:02:26,160
because of decisions made in the past.
It also describes problems that we're creating for 

30
00:02:26,160 --> 00:02:31,520
the future, if we choose limited solutions now.
Google engineers suffered with technical debt for 

31
00:02:31,520 --> 00:02:37,068
several years, as they continued supporting older 
versions of the Android mobile phone operating system. 

32
00:02:37,068 --> 00:02:42,703
In 2019, they stopped supporting 
Android versions from 2011 and earlier. 

33
00:02:42,703 --> 00:02:46,160
The latest of those versions was only 
8 years old, but the difficulty of 

34
00:02:46,160 --> 00:02:50,804
continuing to support it was restricting 
the future development of later versions. 

35
00:02:52,800 --> 00:02:59,120
The adoption of ASCII led to technical debt. With 
a maximum of 127 characters, it wasn't suitable 

36
00:02:59,120 --> 00:03:05,200
for representing accents, nor the different 
scripts – Chinese, Japanese, Cyrillic, Arabic, 

37
00:03:05,200 --> 00:03:10,400
and so on – that are used throughout the world.
But dropping support for ASCII encoded files 

38
00:03:10,400 --> 00:03:16,240
wasn't an option. ASCII was used to represent 
text in about 27 years worth of files. 

39
00:03:16,240 --> 00:03:20,080
Any replacement encoding had to work for 
all the existing ASCII files out there, 

40
00:03:20,080 --> 00:03:23,886
and with all the programs that were 
continuing to create ASCII output. 

41
00:03:26,240 --> 00:03:30,714
Microsoft and IBM each produced their own 
encodings to support various languages. 

42
00:03:30,800 --> 00:03:34,424
Some of their encodings were 
similar, but not exactly the same. 

43
00:03:34,480 --> 00:03:37,760
I'll include a link to the 
Wikipedia Windows code page article, 

44
00:03:37,760 --> 00:03:42,543
in the resources, if you want to read more.
If you try googling for unicode, 

45
00:03:42,543 --> 00:03:47,840
you can spend days following links. It's a complex 
topic, but the good news is that you won't need to 

46
00:03:47,840 --> 00:03:52,623
know more than the brief history in this video, 
and the information in the next few slides. 

47
00:03:52,880 --> 00:03:56,800
Clearly, there was a need for 
a standard encoding. Only now, 

48
00:03:56,800 --> 00:04:01,473
that standard also had to cater for common 
encodings that had sprung up out of necessity. 

49
00:04:03,840 --> 00:04:09,000
The Unicode Standard attempts to provide encodings 
to cater for all the languages in use worldwide. 

50
00:04:09,280 --> 00:04:13,012
Currently, it supports most of 
them, with more being added. 

51
00:04:13,120 --> 00:04:17,464
If you want to know more about it, IBM's 
documentation is a good place to start. 

52
00:04:17,600 --> 00:04:21,000
I'll switch to that page in my browser ...

53
00:04:29,531 --> 00:04:33,567
The contents, on the left, will let you learn more about Unicode.

54
00:04:33,567 --> 00:04:38,385
You'll probably only want to read the first 4 headings – 
once you get to the sections on z/OS,

55
00:04:38,385 --> 00:04:42,560
they're discussing IBM's 
implementation on their z/OS mainframe computers.

56
00:04:42,560 --> 00:04:47,280
The Why the Unicode Standard page is a short 
summary of what we've covered in these slides. 

57
00:04:47,280 --> 00:04:51,000
Have a read through those four 
sections, to understand what Unicode is, 

58
00:04:51,000 --> 00:04:54,899
and why it's important.
Ok, back to the slides. 

59
00:04:55,992 --> 00:05:02,240
Note that those pages were published in 2015. When 
they talk about Unicode "slowly being adopted for 

60
00:05:02,240 --> 00:05:07,760
use in e-mail, too", that was 6 years ago.
Many email clients, including those from 

61
00:05:07,760 --> 00:05:13,607
Google and Microsoft, now either send emails 
encoded as UTF-8, or have an option to do so. 

62
00:05:13,680 --> 00:05:18,720
What might not have been obvious, is that the 
Unicode standard defines several encodings. 

63
00:05:18,720 --> 00:05:22,240
When the IBM documentation referred to 
Unicode being the preferred character 

64
00:05:22,240 --> 00:05:27,185
set for the internet, it's referring to 
Unicode Transformation Format – 8 bit. 

65
00:05:27,280 --> 00:05:32,387
That's more commonly abbreviated to UTF-8, 
and we'll be using that encoding soon. 

66
00:05:34,320 --> 00:05:39,680
Python generally uses UTF-8 by default, 
on many systems. But when you read that, 

67
00:05:39,680 --> 00:05:43,680
in the Python documentation, it 
can give you the wrong impression. 

68
00:05:43,680 --> 00:05:46,800
On Windows, in particular, 
the open method generally 

69
00:05:46,800 --> 00:05:51,840
doesn't default to UTF-8 – which is the 
main reason we've included this video. 

70
00:05:53,840 --> 00:05:59,600
The text files that we include, in the 
resources for the videos, are encoded in UTF-8. 

71
00:05:59,600 --> 00:06:04,897
If you download text files from the internet, 
that's probably the encoding those files will use too.

72
00:06:04,897 --> 00:06:08,716
If you attempt to read a file in 
Python, and get strange results, 

73
00:06:08,716 --> 00:06:13,250
the first thing to check is the file encoding.
We've seen a strange result already, 

74
00:06:13,250 --> 00:06:17,680
when we attempted to read the Jabberwocky file.
If you're using Linux or a Mac, 

75
00:06:17,680 --> 00:06:23,362
it probably looked fine. But on Windows, you may 
have had a strange result at the end of the file. 

76
00:06:23,680 --> 00:06:27,981
I mentioned that at the time, and in the 
next video we'll see what we can do about it. 

77
00:06:29,191 --> 00:06:33,520
See you in the next video.

