1
00:00:00,000 --> 00:00:03,990
Python's regular expression module has a way that you can pre-compile a regular

2
00:00:04,000 --> 00:00:06,990
expression when you're going to be using it over and over again.

3
00:00:07,000 --> 00:00:10,990
This is an efficiency and it also allows you access to some of its more

4
00:00:11,000 --> 00:00:12,990
advanced functionalities.

5
00:00:13,000 --> 00:00:20,990
So let's go ahead and make a working copy of regex.py, call it regex-working.py.

6
00:00:21,000 --> 00:00:24,990
I will open that working copy up and you'll see here that we have a tight little

7
00:00:25,000 --> 00:00:30,990
loop that's using the same pattern over and over again and it's using it to

8
00:00:31,000 --> 00:00:34,990
search for this pattern, in line after line of text.

9
00:00:35,000 --> 00:00:38,990
So we can get some more efficiency out of this loop by pre-compiling the

10
00:00:39,000 --> 00:00:39,990
regular expression.

11
00:00:40,000 --> 00:00:41,990
So up here we open the file.

12
00:00:42,000 --> 00:00:50,990
We can say pattern = re.compile and then we will come down here and we will

13
00:00:51,000 --> 00:00:57,990
cut and paste our pattern and then when we do the search down here, you can say

14
00:00:58,000 --> 00:00:59,990
pattern like that.

15
00:01:00,000 --> 00:01:04,990
And regular expression module will recognize that that's a pattern object.

16
00:01:05,000 --> 00:01:08,990
That's a precompiled regular expression object.

17
00:01:09,000 --> 00:01:10,990
And we will use that, so when we save this and run it.

18
00:01:11,000 --> 00:01:15,990
You see we get the same result, but it's more efficient because we are simply

19
00:01:16,000 --> 00:01:20,990
compiling that regular expression once rather than otherwise the regular

20
00:01:21,000 --> 00:01:24,990
expression module needs to compile the regular expression each time it uses it

21
00:01:25,000 --> 00:01:28,990
and this is in a tight little loop and we want that loop to be fast.

22
00:01:29,000 --> 00:01:32,990
This also gives us the ability to use some of the regular expression module's

23
00:01:33,000 --> 00:01:39,990
other features like for example, you'll notice in the raven text, the very last

24
00:01:40,000 --> 00:01:43,990
use of the word nevermore, it's not capitalized at the beginning.

25
00:01:44,000 --> 00:01:47,990
And so we don't have this line, "Shall be lifted - nevermore!"

26
00:01:48,000 --> 00:01:49,990
at the end of our results.

27
00:01:50,000 --> 00:01:58,990
So we can tell the regular expression model to ignore case, just by putting in re.IGNORECASE.

28
00:01:59,000 --> 00:02:03,990
And in fact you can use just an I, like that. I like to use the whole thing

29
00:02:04,000 --> 00:02:07,990
because it's more explicit and when I come back to look at it years from now,

30
00:02:08,000 --> 00:02:13,990
after maybe I have not used Python for some reason for a while or not used this

31
00:02:14,000 --> 00:02:16,990
particular feature, I will remember what it means.

32
00:02:17,000 --> 00:02:21,990
So it's something you just need to type out once and so it's worth the extra effort.

33
00:02:22,000 --> 00:02:25,990
So if I save this and run it, now we get that last line as well, because we are

34
00:02:26,000 --> 00:02:29,990
telling the regular expression module to ignore case.

35
00:02:30,000 --> 00:02:33,990
So a number of efficiencies are available if we are using pre-complied

36
00:02:34,000 --> 00:02:34,990
regular expressions.

37
00:02:35,000 --> 00:02:45,990
For example, if I want to do a substitution down here, I can say pattern.sub

38
00:02:46,000 --> 00:02:50,990
comma line, like that and now I have the efficiency of having a

39
00:02:51,000 --> 00:02:51,990
pre-compiled pattern.

40
00:02:52,000 --> 00:02:55,990
I can go ahead and use the regular expression module to do the substitution and

41
00:02:56,000 --> 00:03:00,990
I don't necessarily have to get the result out separately and then use the

42
00:03:01,000 --> 00:03:02,990
string module substitution.

43
00:03:03,000 --> 00:03:07,990
So this allows me to use the regular expression engine for more things

44
00:03:08,000 --> 00:03:11,990
without the inefficiencies of having to compile the pattern over and over and over again.

45
00:03:12,000 --> 00:03:16,990
So if I save this and run it, now we have that same result and it's easier to

46
00:03:17,000 --> 00:03:21,990
read and it flows in with our whole pattern if you're using the regular

47
00:03:22,000 --> 00:03:27,990
expression module and we even have that last line with the lowercase nevermore.

48
00:03:28,000 --> 00:03:31,990
It's still working because the pattern was compiled with the IGNORECASE.

49
00:03:32,000 --> 00:03:35,990
So that's just some of the efficiencies that you can get from using pre-compiled

50
00:03:36,000 --> 00:03:37,990
regular expressions in Python.

51
00:03:38,000 --> 00:03:38,990
It's not hard to do.

52
00:03:39,000 --> 00:03:42,990
I recommend that you do it that way if you're going to be using your regular

53
00:03:43,000 --> 00:03:53,000
expressions in a loop or for any purpose that's more than just once or twice.

