1
00:00:05,360 --> 00:00:10,265
In this video, we'll read some data from 
a text file, and try to make sense of it. 

2
00:00:10,320 --> 00:00:14,240
Download the file country_info.txt 
from the resources for this video, 

3
00:00:14,240 --> 00:00:18,560
and put it into your project directory.
Before we can process data, 

4
00:00:18,560 --> 00:00:21,336
we have to understand what it 
is that we're working with. 

5
00:00:21,336 --> 00:00:24,160
One advantage of text files is 
that we can open them in an editor, 

6
00:00:24,160 --> 00:00:30,800
and examine them. Open country_info.txt in your 
IDE, and we'll have a look at what it contains. 

7
00:00:30,800 --> 00:00:34,819
We've pulled together data for most of the 
countries and territories of the world. 

8
00:00:34,880 --> 00:00:39,976
The first line of the file shows the names 
we've used, for each data item we've included. 

9
00:00:39,976 --> 00:00:45,878
For each country, the data shows the name of the 
capital city, the 2-letter and 3-letter country codes, 

10
00:00:45,878 --> 00:00:49,600
the international dialling code used 
to make phone calls to the country, and the 

11
00:00:49,600 --> 00:00:54,979
timezone that the country is in. The last field 
is the name of the currency used in the country. 

12
00:00:55,040 --> 00:00:57,532
We'll talk more about timezones 
in the next section; 

13
00:00:57,532 --> 00:01:00,560
they're not as straightforward
as this data might imply. 

14
00:01:00,560 --> 00:01:05,120
Ok, the most important thing you have 
to understand, when processing data, 

15
00:01:05,120 --> 00:01:10,560
is the structure of the data.
This data is organised in rows – lines of text. 

16
00:01:10,560 --> 00:01:16,758
On each row, we have seven pieces of information, 
separated by a vertical bar, or pipe symbol. 

17
00:01:16,758 --> 00:01:22,720
Each piece of information is commonly referred to 
as a "field". So we have seven fields in each row. 

18
00:01:22,720 --> 00:01:27,693
When dealing with textual data, there has to 
be some way to identify the different fields. 

19
00:01:27,760 --> 00:01:31,680
In this case, the fields are 
separated by a pipe character. 

20
00:01:31,680 --> 00:01:35,527
Fields could also be separated 
by commas, or tab characters. 

21
00:01:35,680 --> 00:01:41,040
Both of those are used in CSV files 
– CSV stands for Comma Separated Values. 

22
00:01:41,040 --> 00:01:44,579
We'll have a look at a CSV 
file, in another example. 

23
00:01:44,640 --> 00:01:48,000
It's essential that each row 
has the same number of fields. 

24
00:01:48,000 --> 00:01:51,440
Have a look at the entry for 
"Antarctica", on line 10. 

25
00:01:51,440 --> 00:01:55,403
Antarctica doesn't have a capital 
city. As far as I'm aware, 

26
00:01:55,403 --> 00:02:00,080
there are no cities in Antarctica .
It's also missing some other information. 

27
00:02:00,080 --> 00:02:05,040
There's no international dialling code, because 
there aren't any telephone exchanges there. 

28
00:02:05,040 --> 00:02:09,287
The currency field is also blank – 
there are no shops in Antarctica either. 

29
00:02:09,360 --> 00:02:11,920
And finally, when you understand timezones, 

30
00:02:11,920 --> 00:02:17,835
you'll appreciate that the polar continents span 
all timezones. So we don't have a timezone field. 

31
00:02:17,920 --> 00:02:22,560
When a piece of data is missing, it's 
essential that the separator is still present. 

32
00:02:22,560 --> 00:02:26,560
We've got seven fields, so each 
row has to have six separators. 

33
00:02:26,560 --> 00:02:31,490
There's no data for four of the fields in 
Antarctica, but we still need the separators. 

34
00:02:31,490 --> 00:02:35,840
If you're not sure why that is, it'll become 
obvious when we start to process this data. 

35
00:02:37,600 --> 00:02:42,080
Making sense of data is called parsing.
When we talk about parsing data – 

36
00:02:42,080 --> 00:02:46,321
or parsing code – it means splitting 
it up into logical components. 

37
00:02:46,400 --> 00:02:50,395
It's something we do all the time, when 
reading or listening to someone speak. 

38
00:02:50,480 --> 00:02:54,819
We do it automatically, so it's probably 
something you've never really thought about. 

39
00:02:54,880 --> 00:02:57,440
But if we're going to get a 
computer to understand the data, 

40
00:02:57,440 --> 00:03:00,438
we have to understand it ourselves, first. 

41
00:03:02,880 --> 00:03:08,204
The Python compiler has to parse our code, before 
it can compile it into something that will run. 

42
00:03:08,204 --> 00:03:14,100
In that sense, our code is data, and we 
send the data (code) to the compiler. 

43
00:03:14,160 --> 00:03:18,942
I'll probably talk about parsing quite often, 
as we look at more advanced features of Python, 

44
00:03:18,942 --> 00:03:25,140
so now you know what it means.
Let's get Python to parse this country data. 

45
00:03:25,140 --> 00:03:31,840
Back in IntelliJ, create 
a new Python file called countries.py. 

46
00:03:35,600 --> 00:03:40,013
In the previous example, we passed the 
filename to open as a string literal. 

47
00:03:40,080 --> 00:03:43,995
It's good practice to store the 
filename in a variable, instead. 

48
00:03:43,995 --> 00:03:48,579
That reduces the risk of errors, if you 
use the same file repeatedly in your code: 

49
00:03:56,080 --> 00:03:59,840
Next, we'll open the file, and 
iterate over the lines of text: 

50
00:04:13,840 --> 00:04:18,399
Ok, how are we going to process this data?
We've seen that each field is separated 

51
00:04:18,399 --> 00:04:23,440
from the next by a pipe character. And we've 
also used Python's split method before. 

52
00:04:23,440 --> 00:04:26,788
We can split the line of text 
up, using the pipe character, 

53
00:04:26,788 --> 00:04:30,960
and get a list containing each of the fields.
That seems like a good approach 

54
00:04:30,960 --> 00:04:32,696
– let's give it a try:

55
00:04:45,907 --> 00:04:47,811
Run the program.

56
00:04:50,160 --> 00:04:53,361
Scroll up through the output, to 
make sure there are no errors. 

57
00:04:53,440 --> 00:04:57,428
That's looking good! We've extracted 
each field from the text file, 

58
00:04:57,428 --> 00:05:02,080
and put them into a list for each country.
There's a minor problem – the last field ends 

59
00:05:02,080 --> 00:05:07,564
with a newline character. We can fix that 
– as we've seen – using the strip method. 

60
00:05:07,760 --> 00:05:11,896
Rather than add another line of code, we 
can strip the line before splitting it: 

61
00:05:26,240 --> 00:05:32,192
Run the program again, and scroll
back up to the start of the output. 

62
00:05:33,920 --> 00:05:39,103
That's removed the newlines.
Note that the list for
"Antarctica" has seven fields, 

63
00:05:39,103 --> 00:05:44,906
with four of them just empty strings.
Ok, we now know that we can read this data successfully. 

64
00:05:44,906 --> 00:05:48,080
The next step is to 
put it into a more useful format. 

65
00:05:48,080 --> 00:05:53,200
A dictionary could be useful, so we'll do that.
I'll see you in the next video.

