1

00:00:00,360  -->  00:00:05,970
Welcome to lesson 27 which roughly covers pages one hundred fifty nine to 163 of the automate the boring

2

00:00:05,970  -->  00:00:08,150
stuff with Python textbook.

3

00:00:08,190  -->  00:00:13,200
So unless lesson you learned about how you can use a carrot to make negative character classes you can

4

00:00:13,200  -->  00:00:18,510
also use the carrot symbol at the start of a regular expression to indicate that the match has to occur

5

00:00:18,510  -->  00:00:24,000
at the beginning of the searched text and likewise you can put a dollar sign at the end of the regular

6

00:00:24,000  -->  00:00:29,730
expression to indicate that the string has to match at the end with this regular expression pattern

7

00:00:29,730  -->  00:00:30,080
.

8

00:00:30,120  -->  00:00:31,320
I'll show you what I mean.

9

00:00:31,320  -->  00:00:37,130
Let's go ahead and import the r e Marshall so that we have access to all the regular expression functions

10

00:00:37,410  -->  00:00:38,700
and I'll create

11

00:00:41,340  -->  00:00:47,310
a regular expression object where it has to begin with the word Helo's call or eadem compile to create

12

00:00:47,310  -->  00:00:53,940
a regular expression object pesudo raw string and we'll have a carrot at the very beginning and then

13

00:00:53,940  -->  00:01:01,320
type hello So means not only does this pattern have to have capital H N and EHLO not only looking for

14

00:01:01,320  -->  00:01:06,030
this text pattern here but this text pattern has to occur at the very beginning of the string or else

15

00:01:06,030  -->  00:01:07,120
it doesn't count.

16

00:01:07,500  -->  00:01:11,550
Hello in the middle of a string that's not considered a match.

17

00:01:12,210  -->  00:01:16,230
So begins with Hello regular expression.

18

00:01:16,290  -->  00:01:21,550
Called the search method and will have it search the text Hello there.

19

00:01:21,870  -->  00:01:26,300
And this indicates that it did find a match it returns a match object.

20

00:01:26,310  -->  00:01:36,150
However if we tried calling search and were looking for the string you said hello this would return

21

00:01:36,420  -->  00:01:42,870
the non-value because even though we have this text appearing here it doesn't begin at the very start

22

00:01:42,900  -->  00:01:46,800
of the string which is what this character character says.

23

00:01:46,800  -->  00:01:54,710
So instead of returning a match object this just returns the non-value.

24

00:01:54,780  -->  00:01:56,280
Similarly we could have something that

25

00:02:01,740  -->  00:02:10,680
has the dollar sign at the end of the string to say that this pattern has to be found at the very end

26

00:02:10,680  -->  00:02:11,480
of the string.

27

00:02:11,640  -->  00:02:16,230
That's what this dollar sign character means.

28

00:02:16,230  -->  00:02:25,710
So if we use this regex object to search the string hello world this would be a match because this world

29

00:02:25,800  -->  00:02:33,780
text because this world pattern is found inside the string and also at the very end of where we had

30

00:02:33,780  -->  00:02:36,460
some other random text there.

31

00:02:36,480  -->  00:02:40,450
Then this search method call would return the non-value.

32

00:02:40,800  -->  00:02:43,600
It didn't find this pattern.

33

00:02:44,150  -->  00:02:50,220
And you can use both of these to indicate that the rejects pattern has to be not only just inside the

34

00:02:50,220  -->  00:02:53,410
text that you're searching but it has to be the entire text.

35

00:02:54,060  -->  00:03:00,710
So let's create a regular expression called all digits all stored in a variable called all digits rejects

36

00:03:00,790  -->  00:03:02,210
.

37

00:03:02,790  -->  00:03:07,790
And so the pattern will started with a carrot symbol so it says the entire text has to begin with this

38

00:03:07,790  -->  00:03:08,500
patter.

39

00:03:08,700  -->  00:03:11,400
We'll just say slash D.

40

00:03:11,820  -->  00:03:12,960
And then one or more.

41

00:03:12,960  -->  00:03:18,070
So the shorthand character class meaning a numeric digit character.

42

00:03:18,090  -->  00:03:23,700
This means that the pattern is one or more of these numeric digit characters and we'll have a dollar

43

00:03:23,700  -->  00:03:29,040
sign at the end of it saying that the string that we're searching also has to end with this pattern

44

00:03:29,140  -->  00:03:30,630
.

45

00:03:31,920  -->  00:03:38,610
So if we tried searching some string like this huge number right here then that would be a match the

46

00:03:38,610  -->  00:03:44,720
entire string has to match because it has both begin and end with this pattern.

47

00:03:45,080  -->  00:03:51,930
Overall we had something like the lowercase letter X that would split this up and that would return

48

00:03:51,930  -->  00:03:54,230
none because it does not match the pattern.

49

00:03:54,240  -->  00:03:58,010
So if you think about it this slash D-plus means one or more digits.

50

00:03:58,080  -->  00:04:04,080
And so you know this string right here it does begin with this pattern because right here it begins

51

00:04:04,080  -->  00:04:08,640
with one or more digits and it also ends with this pattern.

52

00:04:08,640  -->  00:04:15,420
One or more digits but having both of these means that the entire string has to both begin and end with

53

00:04:15,690  -->  00:04:17,830
just this pattern.

54

00:04:18,630  -->  00:04:25,800
So this string does not fit that criteria because it has a non digit character and that's why this returns

55

00:04:25,920  -->  00:04:28,000
None.

56

00:04:28,020  -->  00:04:34,290
Next we'll talk about the wildcard character so well hand character classes like Slash D which stands

57

00:04:34,290  -->  00:04:40,280
for numeric digits or slushed w which stands for word characters like letters or numbers.

58

00:04:40,410  -->  00:04:46,680
Those are nice for having a text pattern that says you know this range of characters but having a period

59

00:04:46,730  -->  00:04:49,400
where a dot in your regular expression syntax.

60

00:04:49,400  -->  00:04:53,760
This stands for any character except for the new line.

61

00:04:54,120  -->  00:04:55,270
So I'll show you what I mean.

62

00:04:55,290  -->  00:05:02,410
Let's create a regular expression object that whole store in this variable at rejects we call our EDA

63

00:05:02,430  -->  00:05:02,930
comp..

64

00:05:02,940  -->  00:05:08,700
Create the expression object and I'll have a dom character meaning that can have any single character

65

00:05:08,700  -->  00:05:09,630
here.

66

00:05:09,900  -->  00:05:11,250
And then the letters 80.

67

00:05:11,310  -->  00:05:20,360
So the pattern is anything followed by 80 like you do a find all on a string that was something like

68

00:05:21,260  -->  00:05:26,570
The Cat In The Hat sat on the flat mat.

69

00:05:27,500  -->  00:05:32,960
And so this is find all there is no groups here so it just returns a list of strings and each of these

70

00:05:32,950  -->  00:05:39,230
strings are a match you see cat here the dot which means anything in this case it's matching a C for

71

00:05:39,230  -->  00:05:41,900
had the character that means anything.

72

00:05:41,900  -->  00:05:44,270
In this case matches in H.

73

00:05:44,320  -->  00:05:45,840
You see that with all of these.

74

00:05:46,040  -->  00:05:52,850
Notice here flat isn't the word that's matched because the DOT character is only looking for a single

75

00:05:53,090  -->  00:05:54,160
character.

76

00:05:54,890  -->  00:05:56,870
And that's why it's just let.

77

00:05:56,880  -->  00:05:58,550
Try changing that.

78

00:05:58,580  -->  00:06:09,230
So just copy this and let's say it's a character and it could be I don't know one or two of those characters

79

00:06:09,230  -->  00:06:09,920
.

80

00:06:09,920  -->  00:06:16,280
So basically the letters 80 that are just preceded by one or two characters of anything.

81

00:06:16,430  -->  00:06:23,000
If we tried calling find all on that well you can see it's probably not what you were expecting because

82

00:06:23,240  -->  00:06:27,690
now it's finding two characters that can be anything and that includes whitespace character.

83

00:06:27,700  -->  00:06:31,280
So we have a space in a C space and an H.

84

00:06:31,280  -->  00:06:34,010
We do have this F and L character.

85

00:06:34,040  -->  00:06:40,520
So when I say it means anything except the new line I do mean any character so common thing that's done

86

00:06:40,520  -->  00:06:44,930
is the pop star patterned in regular expressions.

87

00:06:44,930  -->  00:06:48,790
And so this combination if you remember dot means any character at all.

88

00:06:48,860  -->  00:06:51,350
And the star means zero or more.

89

00:06:51,500  -->  00:06:58,000
So this dot star syntax basically means anything any pattern whatsoever.

90

00:06:58,160  -->  00:07:00,250
And this can be really useful at times.

91

00:07:00,710  -->  00:07:02,030
I'll give you an example.

92

00:07:02,210  -->  00:07:08,120
Let's create name rejects equals Ari compile and.

93

00:07:08,210  -->  00:07:10,820
Well first I'll just copy that for later.

94

00:07:10,820  -->  00:07:14,990
First let's say we had some string value that looks something like this.

95

00:07:15,020  -->  00:07:19,330
First name al last name.

96

00:07:20,590  -->  00:07:22,210
Squired.

97

00:07:22,850  -->  00:07:29,570
So we had text like this and we wanted to pull out the first name and also the last name is me kind

98

00:07:29,570  -->  00:07:29,950
of hard.

99

00:07:29,950  -->  00:07:34,630
I mean I guess we could have some code like.

100

00:07:35,690  -->  00:07:37,070
I don't know.

101

00:07:37,910  -->  00:07:51,200
I find that maybe that first semi-colon would give you the index 10 so maybe plus 2 which would be 12

102

00:07:51,230  -->  00:07:56,900
and then we could use that information to later you know to get a slice.

103

00:07:56,900  -->  00:07:58,690
So it starts off it cuts.

104

00:07:58,790  -->  00:08:03,780
It starts off with the first name and it cuts off this label and we could find some more code like maybe

105

00:08:03,800  -->  00:08:10,520
have find for that last name part and then do some more index math and get the right indexes or we can

106

00:08:10,520  -->  00:08:13,460
use slices to pull out these individual strings right here.

107

00:08:13,460  -->  00:08:15,750
But that's kind of a pain.

108

00:08:16,060  -->  00:08:24,100
Instead let's create a regular expression that does that for us using dot star called precompile.

109

00:08:24,140  -->  00:08:30,950
So in this case what we're looking for is that first name text that's always going to be static right

110

00:08:30,950  -->  00:08:31,260
there.

111

00:08:31,280  -->  00:08:38,720
And then followed by space and then we can create a group that has dot star in it all by another space

112

00:08:38,720  -->  00:08:39,060
.

113

00:08:39,060  -->  00:08:48,310
We didn't have last name followed by a space and then dot star inside of the second group.

114

00:08:49,520  -->  00:08:50,730
So I'm going to copy that.

115

00:08:50,760  -->  00:08:58,940
And now if we call we call find all past that string that we were looking at Rover for find all this

116

00:08:58,940  -->  00:09:04,310
is a regular expression that does contain groups so it's going to return a list of tuples of strings

117

00:09:04,490  -->  00:09:09,380
each tuple are going to be the matches but it's only going to find one match right here and inside that

118

00:09:09,380  -->  00:09:12,760
one tuple are going to be one string for each of the groups.

119

00:09:12,770  -->  00:09:19,480
And so this basically saying whatever he's going to be the string of the first set.

120

00:09:20,010  -->  00:09:21,190
The text of the first ring.

121

00:09:21,200  -->  00:09:24,690
And then this part will be whatever for the last.

122

00:09:24,740  -->  00:09:26,710
The second string.

123

00:09:26,780  -->  00:09:29,910
So what this regular expression object is basically saying is OK.

124

00:09:29,960  -->  00:09:35,750
Look for the text first named colon space and then whatever you find after that.

125

00:09:35,780  -->  00:09:39,440
That'll be the first name going up to this next part.

126

00:09:39,440  -->  00:09:41,420
This last name part in this string.

127

00:09:41,460  -->  00:09:43,170
We're going to be looking at last name Colin.

128

00:09:43,280  -->  00:09:46,400
And then whatever comes after that.

129

00:09:46,520  -->  00:09:51,540
And so this makes it really easy just by specifying the pattern then doing all these calls to find in

130

00:09:51,570  -->  00:09:53,450
all the A list slicing and everything else.

131

00:09:53,450  -->  00:09:58,430
Here you can just do it in two lines of code and you get the text that you want.

132

00:09:58,430  -->  00:10:01,350
And one thing to keep in mind is that Dogstar uses greedy mode.

133

00:10:01,370  -->  00:10:07,400
It will always try to match as much text as possible so use it in a non-greedy fashion.

134

00:10:07,400  -->  00:10:10,550
You want to have dot star question mark.

135

00:10:10,550  -->  00:10:17,700
This is kind of how you this is kind of like putting the question mark after that curly brace syntax

136

00:10:17,810  -->  00:10:27,210
we want to have a non-greedy match but say we're going to have the string angle bracket to serve humans

137

00:10:27,540  -->  00:10:33,640
close angle bracket for dinner close angle bracket.

138

00:10:34,170  -->  00:10:36,560
So this is the text that we're going to start searching.

139

00:10:37,130  -->  00:10:45,690
Most create a non-greedy regular expression and the pattern for this is going to be opening angle bracket

140

00:10:45,810  -->  00:10:49,560
and then dot star.

141

00:10:49,580  -->  00:10:56,460
Make this a group start question mark and enclose angle bracket

142

00:10:57,520  -->  00:11:00,970
.

143

00:11:01,160  -->  00:11:08,340
Now when we call find all for that string that we're going to search this is going to return a string

144

00:11:08,670  -->  00:11:11,220
or a list with a string to serve humans.

145

00:11:11,220  -->  00:11:13,490
That's because it's saying just do a non-greedy match.

146

00:11:13,490  -->  00:11:19,910
We're looking for anything as long as we have the opening angle bracket and then the close bracket but

147

00:11:19,940  -->  00:11:24,800
in between that can be anything but you know as little anything as possible.

148

00:11:24,920  -->  00:11:29,370
So Python will just start saying hey here's the opening angle bracket then we're just going to match

149

00:11:29,430  -->  00:11:31,980
anything until we see a closing bracket.

150

00:11:32,010  -->  00:11:34,780
We're just going to do this as a non-greedy matching.

151

00:11:34,800  -->  00:11:36,280
So here's the first one.

152

00:11:36,330  -->  00:11:37,730
Here's the first close bracket.

153

00:11:37,770  -->  00:11:41,640
We're just going to stop right there because we're not greedy.

154

00:11:41,770  -->  00:11:49,680
If we create a greedy version of this opening angle bracket and then just dot star which we'll do a

155

00:11:49,680  -->  00:11:50,490
greedy match

156

00:11:55,410  -->  00:12:01,640
we search the string now you'll find it actually says hey OK here's that opening angle bracket.

157

00:12:01,650  -->  00:12:08,100
OK so good so far then we'll find anything and then you'll notice hey we can actually match even more

158

00:12:08,100  -->  00:12:08,520
text.

159

00:12:08,520  -->  00:12:14,240
If we go past this first close being close bracket this is also a close bracket.

160

00:12:14,310  -->  00:12:20,760
So we could fulfill we could match that pattern by just matching everything here.

161

00:12:20,760  -->  00:12:25,670
And so that's how you and that's how you get the string to serve humans for dinner.

162

00:12:26,220  -->  00:12:31,650
So before when I was talking about the dot syntax you remember I said that this matches any character

163

00:12:31,680  -->  00:12:34,500
except for the new line character.

164

00:12:34,500  -->  00:12:39,130
Let's say we have a string that's RoboCup three prime directives.

165

00:12:39,660  -->  00:12:43,040
And those are serve the public trust.

166

00:12:43,170  -->  00:12:45,420
We'll have a new line.

167

00:12:45,940  -->  00:12:48,750
Protect the innocent.

168

00:12:49,320  -->  00:12:52,700
And then another new line in the third prime directive.

169

00:12:52,860  -->  00:12:54,720
Uphold the law.

170

00:12:54,720  -->  00:13:00,960
So if we tried printing this out and see these new lines would cause those text to be split across multiple

171

00:13:00,960  -->  00:13:02,150
lines.

172

00:13:02,910  -->  00:13:04,650
So we some dot star.

173

00:13:04,680  -->  00:13:09,820
Regular Expression this is just  dot star.

174

00:13:09,870  -->  00:13:12,570
We'll make this a greedy one.

175

00:13:13,680  -->  00:13:20,090
So even if we search this text with the prime directives on it you'll notice it only matches up to that

176

00:13:20,100  -->  00:13:26,370
first new line because this means ok match whatever character except for New Line and zero or more occurrences

177

00:13:26,370  -->  00:13:27,510
of that.

178

00:13:27,870  -->  00:13:30,360
And since this is a great deal match as much as possible.

179

00:13:30,390  -->  00:13:35,910
So basically ill match until it reaches a new line because the DOT can be any character except for a

180

00:13:35,910  -->  00:13:36,790
new line.

181

00:13:36,810  -->  00:13:42,360
So once it has this newline character right here Bennett says OK that's the first match that we found

182

00:13:43,140  -->  00:13:49,480
but there is a way to get the DOT to mean every character like truly every character including new lines

183

00:13:50,700  -->  00:13:55,050
and that is by passing a second argument to the compile function.

184

00:13:55,280  -->  00:13:56,910
And I'll be the r e.

185

00:13:57,360  -->  00:14:00,210
Dot dot all variable.

186

00:14:00,750  -->  00:14:06,150
So this is just more configuration thing that you can pass to the compile function and it says OK in

187

00:14:06,150  -->  00:14:12,490
this regular expression dots here truly mean everything including new lines.

188

00:14:12,490  -->  00:14:17,880
So now if we try to search that Prime Directive string that dot means everything.

189

00:14:17,880  -->  00:14:20,330
So this truly means match everything.

190

00:14:20,340  -->  00:14:23,720
And also as much as possible because it's a greedy match.

191

00:14:23,820  -->  00:14:29,300
You can see here it now matches the entire string.

192

00:14:29,790  -->  00:14:32,730
So having the second argument to the compile function is pretty useful.

193

00:14:32,730  -->  00:14:37,750
There's also another one where you can have it do a case insensitive regular expression match.

194

00:14:37,800  -->  00:14:44,850
So let's do that Velle regular expression example that we had in the last lesson in creating a regular

195

00:14:44,850  -->  00:14:48,480
expression object for all the vowel characters so e.

196

00:14:48,510  -->  00:14:49,750
Oh and you.

197

00:14:50,010  -->  00:14:53,220
And let's say I forget to add the capital vowels.

198

00:14:53,430  -->  00:15:00,470
Capital AB I O you say I forgot to add that to my character class.

199

00:15:00,630  -->  00:15:04,110
Now if I'm searching a string and something like.

200

00:15:04,110  -->  00:15:14,190
How is your programming book talk about Robocop so much Whoops.

201

00:15:14,370  -->  00:15:21,150
Actually I do find all you see it's returned all the vowels but it's only returning the lowercase files

202

00:15:21,150  -->  00:15:23,270
because that's what I told it to do technically.

203

00:15:23,340  -->  00:15:29,630
So this capital A doesn't actually appear but I can have Python do a case insensitive matching.

204

00:15:29,670  -->  00:15:35,340
I can tell to ignore all casing when I'm creating or regular expression by passing it.

205

00:15:35,350  -->  00:15:36,590
R e.

206

00:15:37,500  -->  00:15:41,040
Ignore case or just for short.

207

00:15:41,130  -->  00:15:42,440
You can also pass r e.

208

00:15:42,510  -->  00:15:48,990
I and this makes it say OK whenever you have a lowercase or uppercase character or whatever it doesn't

209

00:15:48,990  -->  00:15:49,480
matter.

210

00:15:49,500  -->  00:15:51,200
This will match both.

211

00:15:51,300  -->  00:15:58,230
So this means Machon a lowercase a or an upper case a match or lowercase Z or an uppercase e.

212

00:15:58,800  -->  00:16:05,940
Now when we do this same exact code you see the Capitol vowels are also included in the matched text

213

00:16:05,940  -->  00:16:06,610
.

214

00:16:06,660  -->  00:16:12,840
So to recap the carrot character at the beginning of the regular expression string means that the text

215

00:16:12,840  -->  00:16:17,760
it's matching has to begin with this pattern and the dollar sign at the end means that the text has

216

00:16:17,760  -->  00:16:20,530
to end with that pattern.

217

00:16:21,060  -->  00:16:26,010
And if you use both that means the entire string must match the pattern.

218

00:16:26,200  -->  00:16:32,300
The Dutch character is a wildcard character it matches anything except new lines but you can also pass

219

00:16:32,350  -->  00:16:39,330
are the dot dot all as the second argument to compile function to make the dots truly match new lines

220

00:16:39,330  -->  00:16:40,560
as well.

221

00:16:40,620  -->  00:16:47,050
Or you can pass R E Capital I as the second argument to compile and make the matching case insensitive

222

00:16:47,060  -->  00:16:47,670
.

223

00:16:48,030  -->  00:16:54,390
And when you combine the dot character with the asterisk or star character you form the dot star which

224

00:16:54,390  -->  00:17:00,070
is a common way of saying match anything at all because the data is anything and the star means zero

225

00:17:00,100  -->  00:17:02,730
or more of those Anything characters
