1
00:00:01,649 --> 00:00:09,760
So we looked at the sequence matcher
class which returns the ratio method of this

2
00:00:09,760 --> 00:00:17,679
class. It returns this number, but now we
need to see how to get the most similar

3
00:00:17,679 --> 00:00:25,689
word out of a list, out of the keys
of a dictionary, so let's focus first on

4
00:00:25,689 --> 00:00:33,640
on a list on sequence. So a list is
a sequence and from difflib import

5
00:00:33,640 --> 00:00:46,720
get_close_matches. If you do help
get_close_matches you can see the

6
00:00:46,720 --> 00:00:52,900
arguments that you can pass to this
method so the first one in the word so

7
00:00:52,900 --> 00:00:57,250
that will be the word that the
user passes, and then you have a sequence

8
00:00:57,250 --> 00:01:01,960
of possibilities and so word is a
sequence for which close matches are

9
00:01:01,960 --> 00:01:06,429
desired typically a string, so typically
this would be a string that you want to

10
00:01:06,429 --> 00:01:12,520
compare and possibilities is a list of
sequences or a list of strings

11
00:01:12,520 --> 00:01:21,390
you can say. So typically it's a list of
strings. And then you have n equals to 3.

12
00:01:21,390 --> 00:01:27,039
That will define the number of the
matches that you want the get close

13
00:01:27,039 --> 00:01:31,840
matches method to return. Let's say you
pass a list of possibilities with six

14
00:01:31,840 --> 00:01:40,689
items and each item has that ratio so
from 1 to 0, from 0 to 1 and this method

15
00:01:40,689 --> 00:01:45,840
will return the three most similar
matches, the most similar strings to the

16
00:01:45,840 --> 00:01:52,630
word passed in here. You can change that
argument value if you and you have the

17
00:01:52,630 --> 00:02:00,369
cut-off, so that's the ratio here.
By default this will return all

18
00:02:00,369 --> 00:02:06,969
with those items that have a ratio of at
least 0.6 or greater than that

19
00:02:06,969 --> 00:02:11,819
so under that the matches will be not
included.

20
00:02:14,370 --> 00:02:23,350
Q2 end, then help and get let me try close
matches. Let's see what we get.

21
00:02:23,350 --> 00:02:33,250
Let's say rain and then you have a
list in square brackets of course.

22
00:02:33,250 --> 00:02:47,770
Let's say help, pyramid and rain. And I'll
leave the n and the cut off as they are by

23
00:02:47,770 --> 00:02:54,010
default. So in this case you get your
rain the reason you get rain is

24
00:02:54,010 --> 00:03:01,150
because the rain had a cut-off of
greater than 0.6 or similar to that. If

25
00:03:01,150 --> 00:03:05,620
you're interested about the cutoff
you can use the sequence matcher but

26
00:03:05,620 --> 00:03:12,010
we're not interested in that in the
specific value. So rain was the

27
00:03:12,010 --> 00:03:17,020
only one who satisfied those
conditions, so we get rain here as a

28
00:03:17,020 --> 00:03:24,580
list. So okay, we were able to get the close
match out of a list, but how do we get

29
00:03:24,580 --> 00:03:31,030
the close match out of our dictionary
keys? We still have the data variable here

30
00:03:31,030 --> 00:03:40,870
which I loaded previously here, up in
this session. I'll load it from the JSON

31
00:03:40,870 --> 00:03:50,920
file so with a data.keys so there's a
method of dictionaries which is called

32
00:03:50,920 --> 00:03:57,700
keys and with that you can get a list of
all the keys of the dictionary, so we're

33
00:03:57,700 --> 00:04:01,780
talking about only the keys, not the
values of the keys. So these are the

34
00:04:01,780 --> 00:04:07,780
words and each of them has a value
associated with it and basically you can

35
00:04:07,780 --> 00:04:15,070
check the type of data.keys just
see what it is. So it's like a dict_keys

36
00:04:15,070 --> 00:04:23,660
object but it behaves like a list.
And what you can do then is apply the

37
00:04:23,660 --> 00:04:33,800
get matches method so you pass
rain there to data.keys. Let's see

38
00:04:33,800 --> 00:04:42,880
what we get. So we go three matches.
We got rain, train and rainy.

39
00:04:43,430 --> 00:04:52,379
You know if you pass an n of 5 you get
more matches there, but it's important to

40
00:04:52,379 --> 00:05:01,319
note that the list here is ordered, so the
first, the very first word is the one with a

41
00:05:01,319 --> 00:05:08,909
higher ratio of similarity. So we're only
interested in that first word. How to get

42
00:05:08,909 --> 00:05:13,710
that? Well, you can ignore this number
since you only want the first word and

43
00:05:13,710 --> 00:05:17,669
you can pass the zero there because you
know this is a list and the first item

44
00:05:17,669 --> 00:05:24,680
has an index of zero and so you get rain
for the rainn mistyped word.

45
00:05:24,680 --> 00:05:30,180
So basically we have the engine to get
the most similar word out of a

46
00:05:30,180 --> 00:05:36,300
sequence now and all we have to do is
to implement that functionality in our

47
00:05:36,300 --> 00:05:40,979
function in here. So first I would just
like to let the user know through a

48
00:05:40,979 --> 00:05:46,830
message they have input a wrong word.
So I just want to let the user know and

49
00:05:46,830 --> 00:05:51,360
so that they can run a program again so
next time the input a new word for now.

50
00:05:51,360 --> 00:05:55,589
But later we'll do something more
intelligent, so for now just think about

51
00:05:55,589 --> 00:06:01,409
this. Give out a message when you mention
your suggestion with the correct word

52
00:06:01,409 --> 00:06:06,539
that the user may have be, may have
had in mind. So let's do that in the next

53
00:06:06,539 --> 00:06:07,000
lecture, see you.

