1
00:00:00,000 --> 00:00:04,990
Bytes and bytearrays are like tuples and lists except instead of containing

2
00:00:05,000 --> 00:00:09,990
arbitrary objects, bytes and bytearrays contain bytes.

3
00:00:10,000 --> 00:00:12,990
8-bit words of data.

4
00:00:13,000 --> 00:00:18,990
An 8-bit word of data can hold up to 256 different values and this is sometimes

5
00:00:19,000 --> 00:00:20,990
a very convenient thing.

6
00:00:21,000 --> 00:00:25,990
In particular, it's convenient for converting strings and this is where you

7
00:00:26,000 --> 00:00:26,990
will see it used often.

8
00:00:27,000 --> 00:00:30,990
You will see it used for other binary things as well but it's often times used

9
00:00:31,000 --> 00:00:36,990
for converting strings. And we have a great example of this right here.

10
00:00:37,000 --> 00:00:40,990
This is a text file that I created for this purpose and when I created it on

11
00:00:41,000 --> 00:00:45,990
my Mac, it had this lovely little pattern of international characters that

12
00:00:46,000 --> 00:00:46,990
makes a little picture.

13
00:00:47,000 --> 00:00:50,990
It's a little viral thing that had been floating around on Facebook that I got

14
00:00:51,000 --> 00:00:54,990
and I thought it would be great for illustrating this problem, because there are

15
00:00:55,000 --> 00:00:58,990
some circumstances where you cannot display it and it doesn't look right,

16
00:00:59,000 --> 00:01:04,990
or if you try to read it as ASCII data in Python, you will get an exception error.

17
00:01:05,000 --> 00:01:10,990
So I loaded it up on the PC that I am using here and I saw this and I went oh, drat.

18
00:01:11,000 --> 00:01:11,990
It's not pretty.

19
00:01:12,000 --> 00:01:15,990
It doesn't look like it, but in fact this is a great illustration of the problem

20
00:01:16,000 --> 00:01:20,990
because this particular system is not handling the UTF-8 International

21
00:01:21,000 --> 00:01:23,990
Characters properly, whereas my other system was.

22
00:01:24,000 --> 00:01:24,990
They are both running the same software.

23
00:01:25,000 --> 00:01:26,990
They are both running Eclipse.

24
00:01:27,000 --> 00:01:29,990
They are both running the same version of Python and yet here we are trying to

25
00:01:30,000 --> 00:01:34,990
display this file here and it looks like this, whereas on my Mac, it looked

26
00:01:35,000 --> 00:01:37,990
different. And you will see it in a moment, we'll show you here, because we

27
00:01:38,000 --> 00:01:41,990
are going to convert it in a way that it will display here and we're going to

28
00:01:42,000 --> 00:01:42,990
use Python to do this.

29
00:01:43,000 --> 00:01:46,990
So this is what the file looks like here on this PC and if you are using a

30
00:01:47,000 --> 00:01:50,990
different operating system and you actually see the pretty fancy characters,

31
00:01:51,000 --> 00:01:54,990
just shh, don't tell anybody and the example will still work just fine.

32
00:01:55,000 --> 00:02:01,990
I'll start by making a working copy of containers.py and we'll call this

33
00:02:02,000 --> 00:02:08,990
containers-working.py and I'll just close this one and we'll open the working

34
00:02:09,000 --> 00:02:13,990
copy and we are going to start by opening the file. I am going to call this file.

35
00:02:14,000 --> 00:02:23,990
fin open and it's called utf8.txt, open it for read, and we are going to set its

36
00:02:24,000 --> 00:02:33,990
encoding as utf_8 and this is the exact character string that you need to use.

37
00:02:34,000 --> 00:02:38,990
This is meaningful inside of Python and that tells Python that when it's

38
00:02:39,000 --> 00:02:43,990
reading this file, that it needs to read it as UTF-8 and ignore whatever the

39
00:02:44,000 --> 00:02:47,990
default encoding is on your system, which is almost certainly something

40
00:02:48,000 --> 00:02:48,990
different than UTF-8.

41
00:02:49,000 --> 00:02:52,990
UTF-8 is really, really useful encoding.

42
00:02:53,000 --> 00:02:58,990
When the Unicode people came up with Unicode, it's this double wide character set

43
00:02:59,000 --> 00:03:02,990
that doesn't work right in normal ASCII systems where normal 8-bit wide text

44
00:03:03,000 --> 00:03:08,990
context and they tried to get the whole world to adopt it and the whole world didn't adopt it.

45
00:03:09,000 --> 00:03:13,990
So they came up with UTF-8, which is a version of Unicode that works in

46
00:03:14,000 --> 00:03:14,990
an 8-bit encoding scenario.

47
00:03:15,000 --> 00:03:20,990
So the first 127 characters of it works exactly like ASCII does.

48
00:03:21,000 --> 00:03:26,990
So you can set your encodings to UTF-8 safely and it will work just fine with

49
00:03:27,000 --> 00:03:30,990
normal ASCII code and then it has this clever system of setting high bits in

50
00:03:31,000 --> 00:03:34,990
order to tell the system that it needs a couple more bytes to represent a

51
00:03:35,000 --> 00:03:35,990
particular character.

52
00:03:36,000 --> 00:03:39,990
And it all happens kind of transparently behind the scenes if your system is

53
00:03:40,000 --> 00:03:46,990
properly implementing UTF-8. And these days most web browsers do handle UTF-8, just fine.

54
00:03:47,000 --> 00:03:50,990
But a lot of desktop systems don't and this one here that I am working at

55
00:03:51,000 --> 00:03:55,990
obviously doesn't. So we are opening this file as UTF-8 and we are telling that

56
00:03:56,000 --> 00:03:59,990
the encoding is UTF-8 and for it to ignore its default encoding.

57
00:04:00,000 --> 00:04:02,990
I am going to go ahead and open an output file.

58
00:04:03,000 --> 00:04:10,990
I am going to call this utf8.html because we are opening the browser, even

59
00:04:11,000 --> 00:04:14,990
though we are not going to put any actual HTML in it.

60
00:04:15,000 --> 00:04:15,990
And we'll open that for write.

61
00:04:16,000 --> 00:04:21,990
We are going to setup a bytearray, we call it outbytes, initialize the

62
00:04:22,000 --> 00:04:26,990
bytearray, with the bytearray constructor.

63
00:04:27,000 --> 00:04:30,990
And a bytearray is a mutable list of bytes.

64
00:04:31,000 --> 00:04:35,990
So it doesn't hold any other kind of object but bytes and we'll start iterating

65
00:04:36,000 --> 00:04:42,990
through the file for line in file in and then we are going to immediately

66
00:04:43,000 --> 00:04:49,990
iterate through the line for character in line because a string is an iterable

67
00:04:50,000 --> 00:04:54,990
object and we are going to use the ord built in.

68
00:04:55,000 --> 00:05:00,990
if ord of c, and that gives us the integral equivalent of that character.

69
00:05:01,000 --> 00:05:04,990
Is greater than 127.

70
00:05:05,000 --> 00:05:11,990
So there is 128 values in UTF-8 that are just normal ASCII and they are 0 through 127.

71
00:05:12,000 --> 00:05:16,990
So if this one is higher than 127, we are going to do something special with it.

72
00:05:17,000 --> 00:05:20,990
And otherwise, we are just going to append it to outbytes.

73
00:05:21,000 --> 00:05:30,990
We are going to say outbytes.append ord of c, like that. And then if it is greater

74
00:05:31,000 --> 00:05:35,990
than 127, we are going to do this fancy thing here. outbytes +=.

75
00:05:36,000 --> 00:05:42,990
When you use the addition operator on a mutable container type. It has the

76
00:05:43,000 --> 00:05:47,990
same effect as appending, but you can append more than one element at a time this way.

77
00:05:48,000 --> 00:05:53,990
So what I am going to do here is I am going to create a bytes object and bytes

78
00:05:54,000 --> 00:06:00,990
are immutable arrays of bytes and I am going to encode a string.

79
00:06:01,000 --> 00:06:06,990
The constructor of bytes will expect a string within an encoding and so a string

80
00:06:07,000 --> 00:06:11,990
is going to be this XML entity with the ampersand and the pound. If you are

81
00:06:12,000 --> 00:06:16,990
familiar with XML entities, they look kind of like that, where inside of here

82
00:06:17,000 --> 00:06:18,990
you can put a decimal value

83
00:06:19,000 --> 00:06:23,990
that will be interpreted as UTF- 16, which is the normal Unicode.

84
00:06:24,000 --> 00:06:28,990
So in there I am going to have a format and I am going to use this format here,

85
00:06:29,000 --> 00:06:32,990
04decimal. I know this is all looking very complicated.

86
00:06:33,000 --> 00:06:37,990
I told you this line is where all the magic happens. And I am going to use format

87
00:06:38,000 --> 00:06:44,990
ord(c) and then the bytes constructor is going to have an encoding, that

88
00:06:45,000 --> 00:06:50,990
encoding is UTF-8, because we use UTF-8 for everything wherever we can.

89
00:06:51,000 --> 00:06:55,990
So now what we have done is, if the character is outside of the normal ASCII range,

90
00:06:56,000 --> 00:07:01,990
we are going to encode it with this XML entity which can be used in an

91
00:07:02,000 --> 00:07:05,990
HTML context and that will allow us to display our fancy little picture.

92
00:07:06,000 --> 00:07:11,990
Otherwise, if it's not greater than 127, if it's in the normal ASCII range, we

93
00:07:12,000 --> 00:07:12,990
just append it to our outbyte.

94
00:07:13,000 --> 00:07:19,990
So now we have an outbytes bytearray which has all of the characters for our

95
00:07:20,000 --> 00:07:22,990
string and now what we need to do is to turn it in to a string.

96
00:07:23,000 --> 00:07:27,990
We'll call it outstring and we'll use this string constructor and we'll

97
00:07:28,000 --> 00:07:36,990
construct it out of outbytes and guess what? We are going to use encoding = 'utf_8'.

98
00:07:37,000 --> 00:07:42,990
Now all we need to do is to print it to our outfile, print (outstr,file =

99
00:07:43,000 --> 00:07:50,990
fout), and we'll print it also to the screen here so we can see it, and we'll

100
00:07:51,000 --> 00:07:53,990
print the word Done.

101
00:07:54,000 --> 00:07:59,990
So this will read our UTF-8 text from our file that we are not able to read on

102
00:08:00,000 --> 00:08:03,990
this system, go ahead and save this so no catastrophe happens.

103
00:08:04,000 --> 00:08:08,990
This will read our UTF-8 text file and it'll read it with the UTF-8 encoding and

104
00:08:09,000 --> 00:08:15,990
it will write it out to our UTF-8 HTML file, and for the characters that are

105
00:08:16,000 --> 00:08:20,990
outside of the normal ASCII range, it's going to replace them with an XML entity

106
00:08:21,000 --> 00:08:22,990
and that's really all that we are doing here.

107
00:08:23,000 --> 00:08:28,990
So we saved it, we are going to run it, and it looks like I have got a typo some

108
00:08:29,000 --> 00:08:33,990
place here. Yes, right there.

109
00:08:34,000 --> 00:08:38,990
I needed an S. That's all right. Save that and we'll run it and there we

110
00:08:39,000 --> 00:08:40,990
have our fancy string.

111
00:08:41,000 --> 00:08:50,990
So this stuff here got converted to UTF- 16 and these are the Unicode values for

112
00:08:51,000 --> 00:08:55,990
each of those fancy characters and now if we refresh our file system because

113
00:08:56,000 --> 00:09:00,990
Eclipse doesn't like to do that for us and we open this up in the little browser

114
00:09:01,000 --> 00:09:05,990
inside of Eclipse, there is our fancy little picture. And so this is what it

115
00:09:06,000 --> 00:09:06,990
looked like in the text file.

116
00:09:07,000 --> 00:09:11,990
This UTF-8 file has some interesting characters in it and so we weren't able to

117
00:09:12,000 --> 00:09:18,990
see that on this system and by encoding them with the Unicode XML entities,

118
00:09:19,000 --> 00:09:22,990
we are able to see it and there we have it.

119
00:09:23,000 --> 00:09:27,990
So the way that we did this is by using a bytearray.

120
00:09:28,000 --> 00:09:32,990
The beauty of a bytearray is that you can operate on character data because

121
00:09:33,000 --> 00:09:37,990
characters are bytes and a bytearray is mutable, so you can insert things,

122
00:09:38,000 --> 00:09:42,990
you can change it up and all we did here we basically used it as an accumulator.

123
00:09:43,000 --> 00:09:47,990
As we went through the string with the bad data in it, if we found an element

124
00:09:48,000 --> 00:09:51,990
that we needed to operate on, we pushed all of these characters onto the

125
00:09:52,000 --> 00:09:57,990
bytearray, using the bytes constructor and appending them to our outbytes which

126
00:09:58,000 --> 00:10:01,990
is a bytearray. Otherwise we just appended the regular character. If it was

127
00:10:02,000 --> 00:10:04,990
within the range we just appended the regular character.

128
00:10:05,000 --> 00:10:08,990
So these characters here just got appended in the normal way, but these

129
00:10:09,000 --> 00:10:15,990
characters, we ended up using these XML entities which represent the Unicode

130
00:10:16,000 --> 00:10:19,990
characters and we got our little fancy guy to display just the way that we

131
00:10:20,000 --> 00:10:21,990
needed him to display.

132
00:10:22,000 --> 00:10:24,990
So that is a very common use of bytearrays.

133
00:10:25,000 --> 00:10:27,990
Bytearrays are a very effective way to do things like this.

134
00:10:28,000 --> 00:10:32,990
You will see an example very much like this one in our example code later on

135
00:10:33,000 --> 00:10:43,000
in the course.

