1
00:00:00,670 --> 00:00:05,550
I'm glad that you're here and I'm glad that you have made it this far in the cause.

2
00:00:05,590 --> 00:00:09,260
And even if you skipped some sections that's not a problem.

3
00:00:09,300 --> 00:00:14,620
I understand that you may just be interested on doing some stock market analysis.

4
00:00:14,940 --> 00:00:16,080
So that's fine.

5
00:00:16,190 --> 00:00:21,340
But be sure that you know about Pandas and the Bokeh library.

6
00:00:21,660 --> 00:00:28,500
So we use pandas to handle data or whether it is stock market data or whether data or anything we use

7
00:00:28,500 --> 00:00:32,990
that for handling data in Python, analyzing, and grouping and etc.

8
00:00:33,630 --> 00:00:37,660
And also we use a balky library which is a great library for visualizing data.

9
00:00:37,860 --> 00:00:39,970
So that was an intro for this lecture.

10
00:00:40,080 --> 00:00:46,650
Now what you'll get from this lecture is you'll learn how to get stock market data and also other kinds of

11
00:00:46,650 --> 00:00:47,410
data.

12
00:00:47,970 --> 00:00:54,930
And the good thing of this is that you will not download these data from a web site or from your browser,

13
00:00:55,440 --> 00:01:03,790
but you'll get them directly from Python, and the way you do that is by using a library called pandas_datareader.

14
00:01:04,320 --> 00:01:16,170
So you need to install that library and you can install it with pip so install pandas underscore

15
00:01:16,320 --> 00:01:18,410
datareader.

16
00:01:19,410 --> 00:01:26,310
So this used to be actually a module under the Pdas library, but now it's a stand alone library.

17
00:01:26,320 --> 00:01:29,450
so this is how you install it.

18
00:01:30,630 --> 00:01:35,670
I have installed it already so it's says requirements already satisfied, but it shouldn't be a problem

19
00:01:35,670 --> 00:01:37,260
to install it.

20
00:01:37,350 --> 00:01:44,940
So once you do that you can use and of course I suggest you to use the Jupiter Notebook.

21
00:01:46,170 --> 00:01:53,160
So I was in this folder and now if I create a Jupiter file here.

22
00:01:54,450 --> 00:02:00,030
So let's call this stock analysis.

23
00:02:00,890 --> 00:02:02,720
Now I should have that file in here.

24
00:02:02,730 --> 00:02:06,180
So this is the notebook that I created.

25
00:02:06,240 --> 00:02:07,480
Great.

26
00:02:07,500 --> 00:02:12,610
So how do you go about getting these data in Python without even opening the browser?

27
00:02:12,660 --> 00:02:16,830
Well, that's actually fairly easy.

28
00:02:16,830 --> 00:02:21,870
For instance let's try to get recent data from Yahoo.

29
00:02:21,930 --> 00:02:26,000
For some companies that sell stock in there.

30
00:02:26,430 --> 00:02:34,890
So let's say Apple and the first thing you want to do is you want to import the required tool from the

31
00:02:34,890 --> 00:02:42,210
pandas_datareader library and one tool that is used to do to get the stock market data is the

32
00:02:42,210 --> 00:02:42,840
data tool.

33
00:02:42,930 --> 00:02:53,280
So from pandas_datareader import data .Now once you do that ATL 

34
00:02:53,400 --> 00:02:57,140
Enter to execute the line, and once you do that.

35
00:02:57,150 --> 00:03:02,900
you want to use a method called DataReader.

36
00:03:03,780 --> 00:03:05,190
And if you ask for help.

37
00:03:05,190 --> 00:03:08,340
So just a question mark there, execute.

38
00:03:08,930 --> 00:03:14,900
You'll see the docstring, so the definition of this function.

39
00:03:15,000 --> 00:03:21,240
So this imports data from a number of online resources and this library can get data from Yahoo

40
00:03:21,240 --> 00:03:28,790
Finance, and Google Finance, and St. Louis FED or FRED and Kenneth French's data libraries.

41
00:03:28,800 --> 00:03:34,750
So and you can get data from probably all the existing companies out there.

42
00:03:35,050 --> 00:03:44,400
And now, to get the data from these resources, and from different companies what you can do is you want to

43
00:03:44,400 --> 00:03:47,570
pass some parameters to this data_reader function.

44
00:03:47,700 --> 00:03:54,240
Now the first parameter you may want to input there is the name parameter which expects the name

45
00:03:54,240 --> 00:03:54,930
of the company.

46
00:03:54,930 --> 00:04:00,720
Actually it's not the name of the company, of the company but it's a stock symbol.

47
00:04:00,810 --> 00:04:12,510
So stock symbols. If you search for stock symbols and you look at this list of symbols from the New York

48
00:04:12,810 --> 00:04:15,100
Stock Exchange, you'll see

49
00:04:15,300 --> 00:04:25,110
that  you have a big list of companies there, and you've got this code here, or you can even use the search

50
00:04:25,110 --> 00:04:26,830
function of Yahoo.

51
00:04:27,270 --> 00:04:30,070
So these symbols are also called tickers.

52
00:04:30,180 --> 00:04:40,680
So you can enter a ticker here, a company name, so Apple and you see the ticker for Apple is AAPL.

53
00:04:41,010 --> 00:04:48,420
So you put that in there and then you want to specify the data source.

54
00:04:48,420 --> 00:04:53,720
So let's try Yahoo. You should get probably the same data.

55
00:04:53,850 --> 00:04:56,190
So Google gathers data, Yahoo gathers

56
00:04:56,210 --> 00:04:56,890
data

57
00:04:56,890 --> 00:04:58,800
on its sites.

58
00:04:58,890 --> 00:05:00,870
So that's not a big issue.

59
00:05:00,910 --> 00:05:03,740
I haven't tried much the data from Google.

60
00:05:03,760 --> 00:05:09,730
I have worked with Yahoo and I know that the maximum resolution, the temporal resolution that Yahoo

61
00:05:09,730 --> 00:05:11,830
offers is one day.

62
00:05:11,830 --> 00:05:16,180
So you can get a row for each day of the stock market data.

63
00:05:16,570 --> 00:05:22,180
You can also get weekly data. Then if you want you can easily reassemble these data into weekly

64
00:05:22,180 --> 00:05:23,930
or monthly with Pandas.

65
00:05:24,100 --> 00:05:25,630
So Yahoo.

66
00:05:25,640 --> 00:05:32,610
And then you have the start time, and the end time of your data.

67
00:05:32,980 --> 00:05:36,210
So let's define this here.

68
00:05:36,310 --> 00:05:45,860
Let's say the start variable, and that would be datetime.datetime. Actually we need to import the datetime first.

69
00:05:45,890 --> 00:05:48,070
Import datetime.

70
00:05:48,730 --> 00:05:52,630
And then we access the datetime class from the datetime library.

71
00:05:52,630 --> 00:06:01,820
So let's get data from 2016 and let's say March 1st.

72
00:06:01,840 --> 00:06:05,970
So this is the start time, and than the end time would be.

73
00:06:06,430 --> 00:06:13,790
So asyou see I'm using the datetime module to specify the date time of my data so.

74
00:06:13,900 --> 00:06:20,240
And the reason for that, the reason I'm putting these date times here is this the start and the end parameters

75
00:06:20,770 --> 00:06:24,610
they expect from you to enter a datetime data type.

76
00:06:24,940 --> 00:06:26,530
So you cannot just enter a string

77
00:06:26,530 --> 00:06:31,680
there that says 2016 and a coma, or an underscore, or anything.

78
00:06:31,720 --> 00:06:34,980
So you need to be strict there with your format.

79
00:06:35,170 --> 00:06:39,320
So let's say March and let's see 10 here.

80
00:06:39,520 --> 00:06:50,200
So for data for 10 days. Then you need to pass this variables here to these parameters, so starts and ends. That's

81
00:06:50,200 --> 00:06:51,520
it.

82
00:06:51,550 --> 00:06:55,420
So don't confuse the names of the parameters with the names of the variables.

83
00:06:55,660 --> 00:06:57,040
They just come to be the same here.

84
00:06:57,080 --> 00:07:02,180
Now we created these.

85
00:07:02,350 --> 00:07:09,760
And what you want to do now is you want to save the product in a variable.

86
00:07:10,390 --> 00:07:19,810
So I'm choosing to have a df variable, and actually if you check the type of all this expression

87
00:07:22,620 --> 00:07:24,300
well you will first get datetime.

88
00:07:24,300 --> 00:07:28,670
it's not defined because I didn't execute this line where I had it datetime.

89
00:07:28,730 --> 00:07:31,130
So execute that, now execute that again.

90
00:07:31,350 --> 00:07:36,380
And now you see that this object is actually a data frame.

91
00:07:36,390 --> 00:07:45,070
So the pandas_datareader, actually the datareader method of the data class of the pandas_datareader

92
00:07:45,070 --> 00:07:49,540
library produces a data frame, so a pandas data frame.

93
00:07:49,650 --> 00:07:59,400
So that means we can store that data frame in a variable. Remove the brackets and just print out the

94
00:07:59,400 --> 00:08:00,860
data frame here.

95
00:08:01,770 --> 00:08:02,840
And here we go.

96
00:08:02,850 --> 00:08:06,140
Let me close this window.

97
00:08:06,330 --> 00:08:09,470
And here's the data. Now,

98
00:08:09,470 --> 00:08:14,870
don't worry for now if you do not understand these attributes. I'll explain them in the next lecture

99
00:08:15,470 --> 00:08:20,820
because in this lecture I just wanted to know how we get around getting data with Python.

100
00:08:21,260 --> 00:08:24,940
So that is a product you get and you can export that in a csv file.

101
00:08:24,960 --> 00:08:26,690
So that's not a problem.

102
00:08:26,810 --> 00:08:35,240
And if you want to you can get data for a longer longer timeframe so you get a long data frame there.

103
00:08:35,690 --> 00:08:43,100
So from January to March. 1st of January and 10th of March. Of course you can change the company very

104
00:08:43,100 --> 00:08:47,530
easily to anything. Let me just put a random name there.

105
00:08:47,740 --> 00:08:52,230
There are quite a lot of companies out there so this should work.

106
00:08:52,540 --> 00:08:53,650
XRA.

107
00:08:54,140 --> 00:08:57,950
I have no idea what companies is this, but you get the idea.

108
00:08:57,950 --> 00:09:03,320
So it doesn't look as big as Apple, or even on Google.

109
00:09:03,320 --> 00:09:05,870
So you pass GOOG there.

110
00:09:06,420 --> 00:09:13,140
And this is a data for Google, and you can even get that from Google directly.

111
00:09:14,420 --> 00:09:26,450
So forth of January you have 743 for open. Let's see what Yahoo says about Google.

112
00:09:26,450 --> 00:09:27,350
So it says the same thing.

113
00:09:27,350 --> 00:09:32,410
That is about getting stock market data.

114
00:09:32,560 --> 00:09:38,650
Now I can show you how to get other kinds of data as well from the pandas_datareader library.

115
00:09:39,020 --> 00:09:41,680
And I would prefer to point you to the documentation.

116
00:09:41,750 --> 00:09:45,740
I don't want to go through every dataset there.

117
00:09:45,730 --> 00:09:51,740
This is a very clear documentation so I'm sure you'll have no problem going through this.

118
00:09:51,770 --> 00:09:59,330
So here you have the data from Yahoo which we just consumed.

119
00:09:59,900 --> 00:10:02,300
Then you can have data from Google.

120
00:10:02,300 --> 00:10:10,370
So you pass Google there just as I did here. I passed the name of a parameter explicitly there, but you can also

121
00:10:10,730 --> 00:10:13,340
pass them in the correct order like this.

122
00:10:13,340 --> 00:10:19,130
So the company name, the datasource name, start and end, and you'll be fine.

123
00:10:19,130 --> 00:10:21,500
So it is the same thing.

124
00:10:21,680 --> 00:10:25,650
Then you have data from FRED, and Fama/French.

125
00:10:26,000 --> 00:10:30,170
And interestingly you have data from World Bank as well.

126
00:10:30,230 --> 00:10:38,200
So instead of doing from pandas_datareader import data as we did here, you do from pandas_datareader

127
00:10:38,220 --> 00:10:39,160
import

128
00:10:39,380 --> 00:10:48,610
WB so World Bank, and then you search for data sets of GDP per capita and and also other attributes there.

129
00:10:49,390 --> 00:10:55,040
And you should use this download function and that gets you data in a data frame as well.

130
00:10:56,710 --> 00:11:00,610
So that's about World Bank and you have other data sets as well.

131
00:11:00,630 --> 00:11:09,070
So you have OECD statistics there, Eurostats and Edgard index. If you ever need them you know

132
00:11:09,070 --> 00:11:15,270
that they are here in this URL, so in the pandas_datareader documentation.

133
00:11:15,820 --> 00:11:19,600
So that's what I wanted to teach you in this lecture.

134
00:11:19,760 --> 00:11:22,580
I will move on by focusing on the stock market data.

135
00:11:22,590 --> 00:11:27,180
So I'll go quickly through this data explaining them what they mean.

136
00:11:27,190 --> 00:11:30,920
So we've got quite a lot of interesting things to do.

137
00:11:30,970 --> 00:11:31,980
Let's move on then.

