1
00:00:00,030 --> 00:00:08,730
Great, so let's go ahead and build this
chart now using rectangle glyphs. So how

2
00:00:08,730 --> 00:00:13,920
is the rectangle defined? Well when you
use the rect methods from the Bokeh.plotting

3
00:00:13,920 --> 00:00:20,699
interface what you pass to the
rect method are 4 mandatory parameters.

4
00:00:20,699 --> 00:00:29,010
So the first parameter is the
x-coordinate of your rectangles, so if

5
00:00:29,010 --> 00:00:34,860
you have one rectangle in your entire
plot, as the first parameter you'd have

6
00:00:34,860 --> 00:00:41,309
to pass the x-value so which should
happen to be somewhere here. If you have

7
00:00:41,309 --> 00:00:46,200
more so we have let's say 8 here, 8 rectangles
we'd have to pass a list with

8
00:00:46,200 --> 00:00:51,570
8 x values and then you'd have to pass
the y values which as you may

9
00:00:51,570 --> 00:00:57,230
understand is the value of this Y axis,
and then you need to pass the width,

10
00:00:57,230 --> 00:01:03,719
the horizontal width and the height, so that's
how you define your rectangle, so you can

11
00:01:03,719 --> 00:01:07,290
pass lists when you have lots of
rectangles like we have here. You can

12
00:01:07,290 --> 00:01:12,090
pass lists or you can pass data frame
series or data frame columns in other

13
00:01:12,090 --> 00:01:25,290
words, so let's do that. P.rect
and the x-axis is, well, the data

14
00:01:25,290 --> 00:01:32,430
datetime values where close is greater
than open. You know if you start

15
00:01:32,430 --> 00:01:38,790
comparing this rect method to the
quadrant method if you use quadrants and

16
00:01:38,790 --> 00:01:48,200
you pass the date there, your quadrant
would start here at this axis, but

17
00:01:48,200 --> 00:01:54,689
normally the central axis of the
candlestick charts starts at the very day,

18
00:01:54,689 --> 00:01:59,579
then it spans 6 hours left and 6 hours
right, so for quadrants you would

19
00:01:59,579 --> 00:02:06,509
also have to do a trick there, so you'd
have to subtract like 6 hours from

20
00:02:06,509 --> 00:02:13,700
this date here so that you end up with
this kind of position of your quadrant,

21
00:02:13,700 --> 00:02:19,440
but for rectangles we don't need to do
that because we immediately get the

22
00:02:19,440 --> 00:02:26,400
center of the rectangle with this method.
So that the x-axis and then you've got

23
00:02:26,400 --> 00:02:35,190
to pass here the central Y coordinator
of your rectangle and so what we know

24
00:02:35,190 --> 00:02:41,420
about this rectangle is the closing
price here an opening price so that

25
00:02:41,420 --> 00:02:50,330
brings us to the point that the central
axis here would be like cause plus open

26
00:02:50,330 --> 00:03:04,110
divided by 2. So df.open plus give
df.close, so that would need to be

27
00:03:04,110 --> 00:03:13,170
inside brackets here divided by 2 and
that's how we define the y-coordinate.

28
00:03:13,170 --> 00:03:20,120
Now we need to pass the width of the
rectangle so which happens to be always

29
00:03:20,120 --> 00:03:29,760
12 hours. So how do we pass 12 over there?
Well this chart, this plot is designed

30
00:03:29,760 --> 00:03:36,540
to accept milliseconds for values along
the x axis so we have date/time here and

31
00:03:36,540 --> 00:03:41,489
this accepts milliseconds. And that means
we need to convert the 12 hours into

32
00:03:41,489 --> 00:03:47,579
milliseconds and I can either calculate
those seconds and pass the final number

33
00:03:47,579 --> 00:03:53,640
directly here or I could do something
else, you know, if I open this code in a

34
00:03:53,640 --> 00:03:59,100
later time and I have forgotten some
things here so to make it more readable

35
00:03:59,100 --> 00:04:04,260
I'll probably create a variable here and
stole that knowledge in these variable.

36
00:04:04,260 --> 00:04:14,070
So let's say hours 12, so 12 hours and
that would be equal to well 12 hours

37
00:04:14,070 --> 00:04:22,260
times 60 we get minutes, times another 60
we get seconds, and then we need millisecond

38
00:04:22,260 --> 00:04:26,570
so that'll be a thousand
so one second has

39
00:04:26,570 --> 00:04:34,640
thousand milliseconds. Then that means
I can pass hours 12 in this third

40
00:04:34,640 --> 00:04:38,600
parameter which is the width of the
rectangle and then we have the height of

41
00:04:38,600 --> 00:04:45,170
the rectangle which would be the
difference between the opening and

42
00:04:45,170 --> 00:04:56,240
closing price, so df.open minus
df.close, now sometimes you get

43
00:04:56,240 --> 00:05:01,790
negative values here, so you need to think
about when open is higher then close or

44
00:05:01,790 --> 00:05:10,520
vice versa. Or you can just pass abs
here and you will be all right, so

45
00:05:10,520 --> 00:05:15,200
whether open is higher than closed or
the other way around this absolute

46
00:05:15,200 --> 00:05:22,400
method abs makes sure that you always
get positive a value out of this which

47
00:05:22,400 --> 00:05:27,950
will be used as the height of the rectangle.
And now what I could do is I could add

48
00:05:27,950 --> 00:05:35,210
the color that will fill these boxes and
the line around these boxes, but before I

49
00:05:35,210 --> 00:05:40,700
do that I want to make sure I have
everything in place, so I want to output

50
00:05:40,700 --> 00:05:50,690
to what I have so far. Output file. Let's
call this candlestick CS for short dot HTML

51
00:05:50,690 --> 00:06:01,660
and then show, the figure is stored in P
so show P. Execute this and see what we get.

52
00:06:01,660 --> 00:06:10,310
Yeah, not bad.
So let's check one of these, so for the first

53
00:06:10,310 --> 00:06:24,500
of March we have from 704 up to 718.
First of March here, yep. Then day two and

54
00:06:24,500 --> 00:06:32,680
three of March do not exist here, so day
two and three they are not there because

55
00:06:32,680 --> 00:06:40,460
close was less than high for this date
and here we are filtering out

56
00:06:40,460 --> 00:06:44,870
data frame rows through this condition
so that's why we don't have these

57
00:06:44,870 --> 00:06:52,569
candlesticks, and this is weekend so we
don't have data for this,

58
00:06:52,569 --> 00:06:58,380
but we have the ticks
in the x-axis, so…

59
00:06:58,380 --> 00:07:14,730
And let's pass the fill color, so let's
say green and the line color, black, and now

60
00:07:14,730 --> 00:07:21,180
the line here is getting along so what I
could do is I could break this line into

61
00:07:21,180 --> 00:07:25,710
two parts, but make sure if you do the
same, make sure you break a line

62
00:07:25,710 --> 00:07:32,760
after a comma, so you can break it here,
or here, or here, but if you break it

63
00:07:32,760 --> 00:07:42,180
here you'll get an error, so and try that
again, so we have green candlesticks now.

64
00:07:42,180 --> 00:07:56,400
Let's quickly go now ahead and copy
this and paste it there so what I want

65
00:07:56,400 --> 00:08:02,190
to do in this case is I just need to
change this operator here, so in this

66
00:08:02,190 --> 00:08:08,640
case you want to plot the candlesticks
where close is less than open, so this

67
00:08:08,640 --> 00:08:15,930
will be the x-coordinate for the center
of the rectangle and the y-coordinate which

68
00:08:15,930 --> 00:08:23,640
will be the same. And width is 12 hours
so this number here and then the height

69
00:08:23,640 --> 00:08:30,450
the same and the fill color will be red
for these charts, so for this it means

70
00:08:30,450 --> 00:08:37,020
that the day opened with a higher price
and then it closed with a lower price

71
00:08:37,020 --> 00:08:43,970
which means that the
stock price went down throughout the day.

72
00:08:43,970 --> 00:08:52,320
So let's see! This is what we get and we'll
have to remove these horizontal gridlines

73
00:08:52,320 --> 00:08:59,510
later so for now let's focus on the
results. Now this may look good

74
00:09:00,070 --> 00:09:05,470
so your script didn't stop and you may
think that everything is okay, but what

75
00:09:05,470 --> 00:09:10,360
you have to do is whenever you're
building some graphs or some outputs

76
00:09:10,360 --> 00:09:18,030
it's good to get some samples out of
that output and check it with your

77
00:09:18,030 --> 00:09:25,510
source data so in this case for instance
we have this second glyph here which

78
00:09:25,510 --> 00:09:32,530
starts here and it ends here but if you
look and so this says for day two of

79
00:09:32,530 --> 00:09:40,110
March, but if you look here you'll see
that the two starts with an open of 719

80
00:09:40,110 --> 00:09:49,480
and it closes at 718 so with just one
value in difference which this one here

81
00:09:49,480 --> 00:09:56,980
looks more like the second row and
non this, and the reason is that here in the

82
00:09:56,980 --> 00:10:04,000
first rectangle we passed this filter
index, so we got only the days where

83
00:10:04,000 --> 00:10:10,140
closes are greater than open, but then if
you didn't realize that I didn't pass

84
00:10:10,140 --> 00:10:18,690
this filtering in the y-coordinate.
So that has messed things up, so Bokeh

85
00:10:18,690 --> 00:10:26,080
is grabbing the first date time for the
filter date time index and then it gets

86
00:10:26,080 --> 00:10:32,380
the first middle value of clothes and
height and then it gets a second and a

87
00:10:32,380 --> 00:10:38,470
third and then you end up with these
wrong glyphs, so what you have to do is you

88
00:10:38,470 --> 00:10:43,620
need to figure out the open and the
close columns too with this condition.

89
00:10:43,620 --> 00:10:49,270
Now you can do that here, but we'll all
suggest that's the reason that I left

90
00:10:49,270 --> 00:10:54,160
this on filters, so I wanted to get to
the point that is better to actually

91
00:10:54,160 --> 00:11:02,320
write some intermediate code, so what do
I mean by that? Well, for instance in this

92
00:11:02,320 --> 00:11:08,080
case we filter the data frame column on
the fly, so inside the rectangle function.

93
00:11:08,080 --> 00:11:13,270
This is good because you end up with
less lines of code, but you lose some

94
00:11:13,270 --> 00:11:19,120
go over your script so if something
wrong happens there you know and you end

95
00:11:19,120 --> 00:11:24,540
up with these kinds of results. In that
case what I would suggest is actually

96
00:11:24,540 --> 00:11:30,130
create some intermediate results, so for
instance in this case what I could do is

97
00:11:30,130 --> 00:11:38,430
I could store this filtered data time
into a variable. Let's say date increase

98
00:11:38,430 --> 00:11:46,920
so when we have an increase there, that
would be equal to that, and when you have

99
00:11:46,920 --> 00:11:55,540
date decrease, this has to be like this
this time and execute that.

100
00:11:55,540 --> 00:12:02,460
Try them out now.
Date increase. You get this filtered

101
00:12:02,460 --> 00:12:11,710
Pandas series so if you check the type
of this, you'll see that this is a time

102
00:12:11,710 --> 00:12:23,980
series, so similarly you can check your
other variable decrease, and so you get

103
00:12:23,980 --> 00:12:30,220
you get 2nd of March, 3rd of March, so 2nd
of March, 3rd of March where the closing

104
00:12:30,220 --> 00:12:38,020
price was less than the opening price so
these two values here. Then you'd have to

105
00:12:38,020 --> 00:12:48,010
pass these values to their corresponding
places in here, so for the green

106
00:12:48,010 --> 00:12:52,920
boxes you'd have to pass date
increase and then here date decrease.

107
00:12:52,920 --> 00:12:58,500
So passing series.
That's one way to do that.

108
00:12:58,500 --> 00:13:07,330
Another way which I'd like to show you
how to do is to probably create a status

109
00:13:07,330 --> 00:13:13,270
column in your data frame, and then you have
like different strings so you'd have two

110
00:13:13,270 --> 00:13:20,440
strings. An increase text or string if
you like, increase text if the closing

111
00:13:20,440 --> 00:13:24,250
price was higher than the opening price
for that particular row and then

112
00:13:24,250 --> 00:13:30,100
decrease if closing price was slower
than opening price. And then we can

113
00:13:30,100 --> 00:13:35,770
extract the date column out of the data
frame by applying the condition that if

114
00:13:35,770 --> 00:13:40,540
the status column value is equal to
increase, give me the date for that

115
00:13:40,540 --> 00:13:45,190
particular row, and then if it is equal
to decrease then again give me that

116
00:13:45,190 --> 00:13:52,210
particular date for that row. So that's a
more standard way of doing things so I'd

117
00:13:52,210 --> 00:13:59,890
like to do it that way. So let's say df
status so we're creating a new column

118
00:13:59,890 --> 00:14:04,630
om the existing data frame and then we
could use a list comprehension there, so we

119
00:14:04,630 --> 00:14:11,050
could create a list on the fly.
So we'll say something like give me

120
00:14:11,050 --> 00:14:17,890
increase if close is higher than
open or decrease if

121
00:14:17,890 --> 00:14:23,110
close is less than open. And we can
write that statement inside here, or you

122
00:14:23,110 --> 00:14:30,160
can create a function that generates
increase and decrease depending om

123
00:14:30,160 --> 00:14:37,300
these values, on the close and the
open values. So let's say define increase

124
00:14:37,300 --> 00:14:42,250
decrease, let's call it like that.
Here we would have to pass two

125
00:14:42,250 --> 00:14:46,810
parameters, so you'd need to pass the
value for the close call and the value for

126
00:14:46,810 --> 00:14:57,460
the open column, so if C is greater
than O let's say value will be equal to

127
00:14:57,460 --> 00:15:11,950
increase elif C is less than O, value will be
equal to decrease. And maybe we have the

128
00:15:11,950 --> 00:15:17,680
very rare occasion where
value being equal.

129
00:15:17,680 --> 00:15:25,240
So if close is equal to open then we
will have let's say equal. And then what

130
00:15:25,240 --> 00:15:32,710
you need to do is return value so make
sure you are outside the if conditional

131
00:15:32,710 --> 00:15:36,560
here but inside the definition of the
function.

132
00:15:36,560 --> 00:15:43,749
So df equals status and we would have C, O
for C, O in zip df.close, df.open.

133
00:15:55,850 --> 00:16:05,180
So let me explain this. What this is
doing is, it says that for the first value

134
00:16:05,180 --> 00:16:13,220
of the status column assigne the value
or that this function returns, so we're

135
00:16:13,220 --> 00:16:21,860
passing there the C and O variable for
these variables in these columns, so what

136
00:16:21,860 --> 00:16:27,620
this will do it will go through the
next row or the data frame and it will

137
00:16:27,620 --> 00:16:34,730
pick you the close value and the open
value, and it will assign the close value

138
00:16:34,730 --> 00:16:40,100
to the C variable and open value to the
open variable, so this will have values

139
00:16:40,100 --> 00:16:52,430
of around 718 for C and 703 for O.
So these values will be passed to this

140
00:16:52,430 --> 00:17:01,910
function so we would have like if 718
is greater than 700 O, then assign the

141
00:17:01,910 --> 00:17:08,030
value increase to the value variable.
And then at the end of the function you

142
00:17:08,030 --> 00:17:13,030
return increase, so you return the value
which is increasing in this case so this

143
00:17:13,030 --> 00:17:19,100
thing here turns to be the string
increase in the first iteration.

144
00:17:19,100 --> 00:17:23,899
So the list will get the value increase for
the first iteration and then in the next

145
00:17:23,899 --> 00:17:29,240
iteration it will get decrease and so on.
And so you get a list of increase and

146
00:17:29,240 --> 00:17:35,720
decrease strings, and that list then
becomes the status column. So I hope that

147
00:17:35,720 --> 00:17:45,289
is clear. And I'll execute this and print
out the df data frame, and here we go!

148
00:17:45,289 --> 00:17:49,940
So increase for the first decrease,
decrease, and then

149
00:17:49,940 --> 00:17:57,289
the three last ones are decrease, sorry
increase because the closing price is

150
00:17:57,289 --> 00:18:04,700
higher than the opening price, that's it.
And now you go ahead here and you don't

151
00:18:04,700 --> 00:18:10,220
need this anymore.
What you need is the df index of course

152
00:18:10,220 --> 00:18:21,950
so you want the dates where
df.status is equal to increase and again

153
00:18:21,950 --> 00:18:33,320
here we need to delete this.
Df.status equals to decrease in this case.

154
00:18:33,320 --> 00:18:39,019
So that's about the x-axis of the rectangles.
So this one here. Now we need to

155
00:18:39,019 --> 00:18:45,909
define the y-axis which is the middle
price, so open minus close divided by 2.

156
00:18:45,909 --> 00:18:53,600
And as I mentioned earlier it is
good to have intermediate results so for

157
00:18:53,600 --> 00:18:58,480
these two we will do the same thing more
or less. So we need to create another

158
00:18:58,480 --> 00:19:06,799
data frame column. Let's say middle.
So that would be df.open minus df.close

159
00:19:06,799 --> 00:19:13,309
and this goes inside brackets
because we need to divide the entire

160
00:19:13,309 --> 00:19:19,029
expression by two. So if you execute that
and then print out the data frame, you'll

161
00:19:22,610 --> 00:19:28,549
see a middle column there which would be…
Oh yeah, we got some funky numbers there

162
00:19:28,549 --> 00:19:38,120
because I had 2. Sorry!
Add that, so df, yeah, this is the average

163
00:19:42,320 --> 00:19:48,039
Now what we could do is in here we access
the middle column values but again with

164
00:19:53,179 --> 00:20:01,440
a condition and the condition is status
equals to increase.

165
00:20:01,440 --> 00:20:12,830
And the same goes for the other set of
rectangles, but this would be decrease.

166
00:20:12,830 --> 00:20:18,750
So that's great! And we also have the
height here, so we need to filter this

167
00:20:18,750 --> 00:20:25,049
old too because we have two different
sets of rectangles, so again we go here

168
00:20:25,049 --> 00:20:39,570
and say the df height equals to the
absolute value of df.close minus

169
00:20:39,570 --> 00:20:46,590
df.open, so that should do it.
Execute, execute the data frame and we

170
00:20:46,590 --> 00:20:57,750
get the height, so 15 which means 18 here
minus 3 you get the idea is 15. And then

171
00:20:57,750 --> 00:21:07,519
we go our candlestick chart, so we
need to replace this with df.height

172
00:21:07,850 --> 00:21:18,090
with where status is equal to
increase, and same thing for the other set

173
00:21:18,090 --> 00:21:29,970
of rectangles and decrease. Great, execute.
And we have an error. Invalid syntax

174
00:21:29,970 --> 00:21:34,799
there, I seem to have passed and assignment
operator which is different from the

175
00:21:34,799 --> 00:21:42,840
equal operator there. And where is that?
It's here. This one, so one more, execute

176
00:21:42,840 --> 00:21:50,970
again, so here is the updated chart which
looks different from this other one so

177
00:21:50,970 --> 00:21:55,529
now we get this short rectangle there.
Let's check this one as well

178
00:21:55,529 --> 00:22:03,110
so the third. 1st, 2nd, 3rd of March.
This starts somewhere

179
00:22:03,110 --> 00:22:15,130
at 719, 712 so third, yeah it's summer at 718
712, so it's looking good. And I'd like to

180
00:22:15,130 --> 00:22:20,260
stop the video here and in the next
lecture we'll go ahead and make this

181
00:22:20,260 --> 00:22:24,880
plot area more visually appealing.
We'll remove these gridlines and also add

182
00:22:24,880 --> 00:22:28,000
the segments, so I'll see you later.

