1

00:00:00,210  -->  00:00:06,030
Welcome to lesson 29 in this lesson we're going to create a program using all of our regular expression

2

00:00:06,060  -->  00:00:07,690
and previous programming.

3

00:00:08,070  -->  00:00:14,340
So say that your boss comes to you with a giant PDA file full of phone numbers and e-mail addresses

4

00:00:14,520  -->  00:00:22,020
and says I need you to copy and paste out every single phone number in this document and this document

5

00:00:22,080  -->  00:00:24,740
is dozens of pages long.

6

00:00:25,080  -->  00:00:31,050
So if you had to do this on your own just copying and pasting each of these lines over and over and

7

00:00:31,050  -->  00:00:37,890
over again this would take you hours especially if you had multiple documents like this instead let's

8

00:00:37,890  -->  00:00:41,930
write a Python program that will do this for us.

9

00:00:42,270  -->  00:00:48,060
Now it'll be a lot easier if we can just go to these documents press control a to select all and then

10

00:00:48,060  -->  00:00:56,780
Control C to copy it and then have our program just read this text off of the clipboard.

11

00:00:57,420  -->  00:01:00,630
So we'll be using the paperclip module for this program.

12

00:01:00,960  -->  00:01:06,450
So in a brand new editor window let's just save this as phone and emailed out pi since this will be

13

00:01:06,450  -->  00:01:08,250
a program that we run over and over again.

14

00:01:08,280  -->  00:01:13,560
We're going to started off with a shebang line which tells Python which version of Python that we want

15

00:01:13,560  -->  00:01:15,200
to run this script.

16

00:01:15,240  -->  00:01:21,390
So I'm on Windows so for me it'll look like Python 3 And next let's create a series of two do comments

17

00:01:21,540  -->  00:01:25,370
that will just sort of create a skeleton of what we want our program to do.

18

00:01:25,530  -->  00:01:31,670
So first we'll have to create a rejects object for phone numbers.

19

00:01:32,230  -->  00:01:40,110
Then we have to create a regexp object for email addresses.

20

00:01:40,120  -->  00:01:43,860
Next we'll have to get the text off the clipboard

21

00:01:46,930  -->  00:01:53,190
and we have to extract the e-mail addresses and phone numbers from this text

22

00:01:56,090  -->  00:02:03,400
and we'll have to copy the extracted phone numbers and e-mails to the clipboard.

23

00:02:03,980  -->  00:02:06,010
So let's do this one at a time.

24

00:02:06,170  -->  00:02:11,910
At the very top of our program underneath the shebang line we're going to have to figure out which modules

25

00:02:11,910  -->  00:02:13,100
we want to import.

26

00:02:13,230  -->  00:02:17,520
So since we'll be dealing with regular expressions we'll have to import the r e module.

27

00:02:17,760  -->  00:02:25,080
And also since we'll be copying and pasting text from the clipboard we have to import the paperclip

28

00:02:25,080  -->  00:02:26,610
module as well.

29

00:02:26,610  -->  00:02:31,410
Now this paperclip model doesn't come with Python so you'll have to install it separately and there

30

00:02:31,410  -->  00:02:34,500
are instructions on doing this in the course notes.

31

00:02:34,560  -->  00:02:38,840
You'll just have to use the pip program to install it.

32

00:02:39,060  -->  00:02:46,710
So the regular expression objects are created by calling eat up compile and we just pass it a string

33

00:02:46,710  -->  00:02:46,970
.

34

00:02:46,980  -->  00:02:51,990
It's usually not helpful to use a raw string because we have a lot of backslashes in these strings usually

35

00:02:51,990  -->  00:02:53,160
.

36

00:02:53,160  -->  00:02:58,650
And I want to use verbose mode for this to pass already verbose as the second argument.

37

00:02:58,800  -->  00:03:06,900
So allow me to use a triple quoted multi-line string and also all this whitespace and even the comments

38

00:03:06,900  -->  00:03:12,960
that I add inside of the regular expression string will be a part of the actual pattern that it needs

39

00:03:12,960  -->  00:03:13,910
to match.

40

00:03:13,920  -->  00:03:16,040
This will make it a lot more readable.

41

00:03:16,200  -->  00:03:21,550
So when we think about the types of phone numbers that I want to be able to collect.

42

00:03:21,550  -->  00:03:29,240
So the really basic type is just a three digit area code with the phone number separated by dashes.

43

00:03:29,250  -->  00:03:31,530
But that's not the only way that it could look.

44

00:03:31,530  -->  00:03:37,590
In fact the area code could be completely optional and not even there or the area code could be surround

45

00:03:37,590  -->  00:03:38,540
by parentheses.

46

00:03:38,580  -->  00:03:42,600
And in so that first Dasch just have a space instead.

47

00:03:43,170  -->  00:03:48,540
And then we could also have phone numbers that have extensions after them so it could look like the

48

00:03:48,540  -->  00:03:56,580
word XTi followed by a number of digits let's say this will be anywhere between 2 and 5 digits long

49

00:03:56,580  -->  00:04:07,590
for the extension but it could look like the x t dot where even just the letter x.

50

00:04:08,640  -->  00:04:12,440
So we just write out some more comments inside this regular expression string.

51

00:04:12,480  -->  00:04:16,580
This will sort of be the skeleton of all the different parts of the regular expression.

52

00:04:16,580  -->  00:04:17,110
I want to make.

53

00:04:17,120  -->  00:04:21,560
So we'll have this area code which is optional.

54

00:04:22,410  -->  00:04:32,130
Now we'll have that first separator comes after the area code won't have the first three digits followed

55

00:04:32,130  -->  00:04:40,470
by another separator followed by the last four digits and then followed by an extension which will also

56

00:04:40,470  -->  00:04:43,920
be optional.

57

00:04:45,530  -->  00:04:50,730
May just have some spaces here by highlighting them all and adding tabs.

58

00:04:50,870  -->  00:04:53,180
OK so the area code is pretty simple.

59

00:04:53,400  -->  00:05:00,630
Could just be three digits or it could be the print to see three digits round by parentheses so I'm

60

00:05:00,630  -->  00:05:06,850
going to put this in a group we use parentheses we'll create this as its own separate group.

61

00:05:07,020  -->  00:05:14,460
So say I'm looking for this group or which I use the vertical pipe character to me or I'm looking for

62

00:05:14,480  -->  00:05:20,070
another group and this one has literal parentheses song you have to escape them with the backslash.

63

00:05:20,130  -->  00:05:26,150
So in opening parentheses followed by three digits followed by a literal closing the C.

64

00:05:26,500  -->  00:05:33,310
And also this entire area code could be optional so I'm going to put it in its own group and this group

65

00:05:33,310  -->  00:05:38,930
means you know either this pattern or this pattern and that group itself will be optional.

66

00:05:38,940  -->  00:05:45,310
Put the question mark right here that says this entire group can appear in the pattern 0 or 1 times

67

00:05:45,370  -->  00:05:45,680
.

68

00:05:46,010  -->  00:05:47,880
And after that we'll just have the separator.

69

00:05:47,880  -->  00:05:50,780
Could either be a whitespace character like this space.

70

00:05:50,860  -->  00:05:56,590
So just use the space shorthand character class or it could be a dash.

71

00:05:56,940  -->  00:05:59,040
I'll just put this in its own group as well.

72

00:05:59,190  -->  00:06:00,600
And next we'll have the first three digits.

73

00:06:00,610  -->  00:06:01,940
This is pretty simple.

74

00:06:02,060  -->  00:06:06,180
Let's have three digit characters that we're looking for followed by another separator.

75

00:06:06,210  -->  00:06:14,350
We'll just say that this will be the dash in between the first three and last four characters or digits

76

00:06:14,830  -->  00:06:19,650
behind the last four digits that we're looking for and then we'll have an extension.

77

00:06:19,650  -->  00:06:23,110
So this is going to be a little complicated let's think.

78

00:06:23,250  -->  00:06:29,630
Could either be the word XTi followed optionally by a literal parentheses.

79

00:06:29,680  -->  00:06:31,140
So I'm sorry.

80

00:06:31,140  -->  00:06:37,660
A literal period would put that literal period that's been escaped with the backslash inside its own

81

00:06:37,650  -->  00:06:41,180
group and then put a question mark after that saying that this is optional.

82

00:06:41,250  -->  00:06:43,380
It could appear or it could not appear.

83

00:06:43,380  -->  00:06:49,030
And then we'll have a space character that follows it followed by a certain number of digits we'll say

84

00:06:49,190  -->  00:06:52,190
so we'll have a certain number of digits here.

85

00:06:52,240  -->  00:06:58,150
We put this in its own group we'll say it could be two to five digits so we'll use that curly brace

86

00:06:58,620  -->  00:07:04,310
meaning that this pattern here for a single digit could appear two to five times.

87

00:07:04,380  -->  00:07:07,980
And of course let's put all of this extension stuff.

88

00:07:08,280  -->  00:07:12,830
Oh whoops almost forgot about this format.

89

00:07:13,780  -->  00:07:15,860
Let's put all of this inside of a group.

90

00:07:16,000  -->  00:07:21,080
So the extension could look like that are actually going to keep going.

91

00:07:21,100  -->  00:07:23,460
Just separate this part out.

92

00:07:23,470  -->  00:07:30,000
Extension could look like this with x t followed optionally by a print of C or followed optionally by

93

00:07:30,000  -->  00:07:32,790
a period and then a space.

94

00:07:32,820  -->  00:07:35,740
Or it could just look like the letter x by itself.

95

00:07:35,910  -->  00:07:40,530
I'll just put this as that word party the extension

96

00:07:45,040  -->  00:07:55,140
of the C extension word part and I'll just cut and paste this on the next line.

97

00:07:56,250  -->  00:08:03,680
So the the extension the number part is optional.

98

00:08:04,060  -->  00:08:09,150
I don't need to put both of these parts inside of a group and make that group optional so you'll have

99

00:08:09,150  -->  00:08:18,160
a opening parentheses here and then a close friend to see here and follow up with a question mark so

100

00:08:18,150  -->  00:08:22,230
that this entire group will be optional.

101

00:08:22,870  -->  00:08:24,990
So this looks pretty complicated right.

102

00:08:25,200  -->  00:08:31,630
But verbose mode allows us to add these comments and also these new lines and other space characters

103

00:08:31,870  -->  00:08:35,360
as part of the regular expression string without changing what the pattern is.

104

00:08:35,470  -->  00:08:40,120
So if we didn't have verbose mode we'd have to use a single line string that would look like this

105

00:08:40,120  -->  00:08:49,740
.

106

00:08:49,810  -->  00:08:55,840
So this line of code is much harder to read and make sense of than this line of code at least we have

107

00:08:55,840  -->  00:08:58,920
the comments here telling us what each of these parts are.

108

00:08:58,920  -->  00:09:04,530
So if we have to go back in and make changes those could later be a lot easier to figure out what we

109

00:09:04,530  -->  00:09:10,560
meant by all this code than if we just had this giant wall of text right here.

110

00:09:10,570  -->  00:09:15,120
So now that I'm done with this part of creating a regular expression for phone numbers I'll get rid

111

00:09:15,120  -->  00:09:18,640
of this to do and I'll just leave this comment here.

112

00:09:18,630  -->  00:09:21,260
This will describe what this part of the code is doing.

113

00:09:21,390  -->  00:09:25,840
And of course compile will return a regular expression object so I'll have to save that to a variable

114

00:09:26,030  -->  00:09:33,670
with just say don't read Jack's next let's move on to create a regular expression for email addresses

115

00:09:34,050  -->  00:09:43,510
to have something like phone rejects equals a compile have a multi-line string and use verbose mode

116

00:09:43,510  -->  00:09:46,090
for this compile function.

117

00:09:46,090  -->  00:09:52,180
So the actual regular expression for e-mail addresses is really crazy because email addresses can have

118

00:09:52,180  -->  00:09:58,090
all sorts of things we're used to seeing them as something at something dotcom but this name part can

119

00:09:58,090  -->  00:10:04,480
also have periods and it can have plus signs in it I think can even have percent and question marks

120

00:10:04,480  -->  00:10:10,740
and it was just baldest handle just dots and plus signs.

121

00:10:10,930  -->  00:10:15,190
I also have underscores as well those show up in email addresses too.

122

00:10:15,370  -->  00:10:23,140
This could be something like edu or dot gov or dot net news domain name right here could also be using

123

00:10:23,140  -->  00:10:25,410
tons of weird characters in it.

124

00:10:25,720  -->  00:10:31,180
So we'll just use sort of the same pattern for this as we do for this.

125

00:10:31,180  -->  00:10:34,840
So let's just use common Skrill a skeleton.

126

00:10:34,900  -->  00:10:44,930
First of all be of the name part and there will be the at symbol followed by the domain name part.

127

00:10:47,110  -->  00:10:51,910
So for the name part we can't just use slash W because we also need to include characters like this

128

00:10:51,910  -->  00:10:57,730
dot and the plus sign and slushed W is a character class just for letters numbers and the underscore

129

00:10:57,730  -->  00:10:58,410
.

130

00:10:58,420  -->  00:11:02,820
So instead we'll have to create our own character class using the square brackets.

131

00:11:02,980  -->  00:11:09,470
So the character class will be able to match lower case letters and also upper case letters and the

132

00:11:09,490  -->  00:11:11,200
digits 0 through 9.

133

00:11:11,230  -->  00:11:17,770
You also want it to match a underscore or a period character or a plus character remember inside of

134

00:11:17,770  -->  00:11:19,930
a character class in between the square brackets.

135

00:11:19,930  -->  00:11:25,410
We don't have to escape these dots and plus signs with backslashes.

136

00:11:25,530  -->  00:11:31,690
That's only something that we have to do outside of the square brackets in a character class like we

137

00:11:31,690  -->  00:11:37,500
did here with the parentheses the literal parentheses that we wanted to match in a pattern.

138

00:11:37,510  -->  00:11:43,540
This is the range of characters that can be in this name part of the e-mail address and we'll be searching

139

00:11:43,540  -->  00:11:47,030
for one or more of them so put a plus sign after that.

140

00:11:47,050  -->  00:11:48,370
Next is the at symbol part.

141

00:11:48,370  -->  00:11:54,790
This is really easy to say at symbol and then the domain name part we'll just go ahead and copy and

142

00:11:54,790  -->  00:11:57,970
paste this part and use that

143

00:12:03,160  -->  00:12:11,890
to get the text off of the clipboard so here we can use the paperclips paperclip module's paste function

144

00:12:12,140  -->  00:12:17,620
and that will return the string of the text that's currently on the clipboard so as just save that into

145

00:12:17,620  -->  00:12:19,510
text.

146

00:12:19,600  -->  00:12:25,180
Next we'll extract the e-mail and phone numbers off of the text.

147

00:12:25,300  -->  00:12:29,530
So here we'll use this phone rejects object.

148

00:12:29,530  -->  00:12:36,550
It has a find all method we'll just pass it that tests to check for this phone regular expression pattern

149

00:12:36,970  -->  00:12:43,010
and find or will return a list of strings for us with each string being a matched phone number.

150

00:12:43,060  -->  00:12:51,040
So I'll just save that in extracted phone as a variable will do the same thing for all the e-mail addresses

151

00:12:51,040  -->  00:12:55,500
as well we'll just use the email regular expression object here.

152

00:12:55,930  -->  00:12:59,820
See that in an in a variable named extracted e-mail.

153

00:13:00,490  -->  00:13:04,570
So we've added a lot of code but we're not really sure if all of this works or not.

154

00:13:04,630  -->  00:13:08,830
So let's just temporarily add some print function calls and just print out

155

00:13:11,830  -->  00:13:14,730
what the values are inside these two variables.

156

00:13:14,830  -->  00:13:22,930
So I'm going to go and instead of selecting all with control a I'm just going to grab some of the text

157

00:13:22,930  -->  00:13:23,330
here.

158

00:13:23,380  -->  00:13:29,150
Let's just say let's highlight just this amount and press control-C to copy that to the clipboard and

159

00:13:29,150  -->  00:13:32,390
then run this program.

160

00:13:32,650  -->  00:13:35,770
So the extracted e-mails seems to work.

161

00:13:35,770  -->  00:13:38,300
This is just a list of email address strings.

162

00:13:38,320  -->  00:13:39,900
That's exactly what we expected.

163

00:13:40,090  -->  00:13:42,970
But this looks a little strange.

164

00:13:43,120  -->  00:13:48,370
It's not a list of strings but actually a list of tuples and inside the tuples.

165

00:13:48,370  -->  00:13:50,060
There are multiple strings.

166

00:13:50,420  -->  00:13:51,700
Oh yes I remember now.

167

00:13:51,760  -->  00:13:57,130
So find all returns something slightly different depending on if there is more than one group in the

168

00:13:57,160  -->  00:14:01,270
regular expression or not e-mail the e-mail rejects has no groups.

169

00:14:01,360  -->  00:14:05,890
So this is going to just return a simple list of strings.

170

00:14:05,890  -->  00:14:11,530
But since there are groups here find all is going to return a list of tuples for each match and each

171

00:14:11,530  -->  00:14:17,030
tuple has several strings in it one string for each of these groups.

172

00:14:17,200  -->  00:14:23,300
And the really simple way to solve that problem is just put everything inside one large group.

173

00:14:23,650  -->  00:14:25,510
And that way group 0.

174

00:14:25,510  -->  00:14:30,760
The very first group is going to cover the entire matched text.

175

00:14:30,760  -->  00:14:37,300
So I'm going to press control as to say and let's run this again so that's good.

176

00:14:37,300  -->  00:14:42,310
Now we can just go through this list and we know that the first string in each of these tuples will

177

00:14:42,310  -->  00:14:46,450
be the full phone number.

178

00:14:46,840  -->  00:14:55,000
So I have to do something like when we all just create a blank list hold all phone numbers start off

179

00:14:55,000  -->  00:15:01,750
as a blank list and then I'll just loop over all of the tuples inside this a list of tuples and extracted

180

00:15:01,750  -->  00:15:08,200
phone I'll just say for phone number in extracted phone.

181

00:15:08,920  -->  00:15:14,470
So on each iteration through this loop the phone number variable will be assigned a single tuple from

182

00:15:14,470  -->  00:15:16,950
this list of tuples in extracted phone.

183

00:15:17,230  -->  00:15:25,120
And I'll just append that phone number to all phone numbers so all the phone numbers append phone number

184

00:15:26,680  -->  00:15:27,870
phone number zero.

185

00:15:27,880  -->  00:15:31,240
I just want that first string in this tuple.

186

00:15:31,300  -->  00:15:38,530
So by the time this loop completes the all phone numbers list will have all of those first strings from

187

00:15:38,530  -->  00:15:44,350
this tuple that was taken from this extracted phone which was a list of tuples.

188

00:15:45,460  -->  00:15:52,900
So basically around the first iteration of that for loop phone number will be set phone number will

189

00:15:52,900  -->  00:16:00,130
be set to this value this tuple and phone number 0 will just be this first string and this will be appended

190

00:16:00,130  -->  00:16:02,000
to the all phone numbers list.

191

00:16:02,030  -->  00:16:04,720
It was check this out right now.

192

00:16:04,830  -->  00:16:09,420
It's extracted phone I'll just print out this all phone numbers list and we can test that.

193

00:16:09,520  -->  00:16:10,570
We've already tested this.

194

00:16:10,570  -->  00:16:17,110
I'll just get rid of that to go back here press control-C to copy and then run this program tester again

195

00:16:17,120  -->  00:16:17,360
.

196

00:16:17,560  -->  00:16:17,890
OK.

197

00:16:17,890  -->  00:16:22,060
That seems to work just fine.

198

00:16:22,090  -->  00:16:28,580
All right so we have this list value with all the phone numbers and also a list of all the e-mail addresses

199

00:16:28,600  -->  00:16:28,780
.

200

00:16:29,080  -->  00:16:34,300
But we don't want to copy this text to the clipboard we don't want all these quotes and commas and everything

201

00:16:34,300  -->  00:16:37,840
would be much nicer if we could just have one phone number per line.

202

00:16:37,840  -->  00:16:41,900
So you just have this phone number on a line and then followed by this phone number on it's own line

203

00:16:41,910  -->  00:16:42,830
.

204

00:16:43,420  -->  00:16:46,000
So we can do that with the join.

205

00:16:46,510  -->  00:16:53,050
Remember join takes a list of strings such as are all phone numbers list and it joins them together

206

00:16:53,050  -->  00:16:58,330
into a single string and this string will be in between each of the strings in this list.

207

00:16:58,330  -->  00:17:04,000
So I'll just put a new line character so there is a newline character in between all of the string the

208

00:17:04,000  -->  00:17:06,420
phone numbers in this list of phone number.

209

00:17:06,460  -->  00:17:17,140
So that'll put them one phone number per line and I'll do the same for the e-mail address as well.

210

00:17:17,140  -->  00:17:21,130
These are just expressions that evaluate to a string that you need to do something with that string

211

00:17:21,130  -->  00:17:21,440
value.

212

00:17:21,460  -->  00:17:29,770
So just stored in a variable to say results equals this string and I'll concatenate a new line to that

213

00:17:30,250  -->  00:17:34,090
and then concatenate all of those e-mails.

214

00:17:34,090  -->  00:17:41,580
So this forms one giant string that is stored in the results variable.

215

00:17:41,610  -->  00:17:49,720
And then I'll just use paper clips copy function to copy that text to the clipboard.

216

00:17:51,390  -->  00:17:51,800
All right.

217

00:17:51,820  -->  00:17:58,240
Let's do a test run of this girl this document this huge hundred page document full of phone numbers

218

00:17:58,240  -->  00:18:02,290
and e-mail addresses that would take me hours to copy out by hand.

219

00:18:02,410  -->  00:18:09,640
And instead of just press control a to select all Control-C to copy this to the clipboard and then just

220

00:18:09,640  -->  00:18:13,540
run my program and there it has no help whatsoever.

221

00:18:13,540  -->  00:18:15,500
There's no print function calls here.

222

00:18:15,700  -->  00:18:21,520
But if this is worked correctly all of those phone numbers and email addresses have been extracted put

223

00:18:21,520  -->  00:18:28,870
into a nicer format and then that text has been copied to the clipboard so it could go to a word processor

224

00:18:28,870  -->  00:18:34,450
program and open up a new document and just paste all of that data to this document and it's a much

225

00:18:34,450  -->  00:18:40,030
nicer format or if I wanted to email this to somebody I could just paste it into the email or I could

226

00:18:40,150  -->  00:18:42,120
paste it into a spreadsheet program.

227

00:18:42,160  -->  00:18:46,670
It's in a format that works a lot nicer if I wanted it in a slightly different format.

228

00:18:46,780  -->  00:18:49,720
I could just change this code right here.

229

00:18:50,060  -->  00:18:54,230
Maybe have I don't know commas in between Instead of new lines.

230

00:18:54,230  -->  00:18:55,670
Basically whatever I it.

231

00:18:56,020  -->  00:19:01,690
And since I have this in it's own program even if I had multiple documents like this I could just open

232

00:19:01,690  -->  00:19:08,050
up the next document press control a press control-C to copy and then run the program again for that

233

00:19:11,500  -->  00:19:17,710
and then just take that data that's now on the clipboard and also paste it to the document as well.

234

00:19:17,730  -->  00:19:25,410
And so that's how writing a very small python program this is just what about 30 lines long.

235

00:19:25,420  -->  00:19:29,860
Just having some programming knowledge can save you hours of time on boring tasks
