1
00:00:05,200 --> 00:00:07,760
In the last video, we saw a situation where using 

2
00:00:07,760 --> 00:00:11,261
discard was a more appropriate 
way to delete items from a set. 

3
00:00:11,261 --> 00:00:15,360
In this video, we'll look at a situation 
where remove is more appropriate. 

4
00:00:15,360 --> 00:00:19,137
I'm going to use the example that I 
mentioned, in the video before last. 

5
00:00:19,360 --> 00:00:23,600
We'll consider a clinical trial, to test 
replacing the anti-coagulant drug Warfarin 

6
00:00:23,600 --> 00:00:29,760
with a different blood thinning agent, Edoxaban.
Once again, we've got some data to work with. 

7
00:00:29,760 --> 00:00:36,121
Create a new Python file called prescription_data, 
and paste in the code from prescription_data.txt, 

8
00:00:36,121 --> 00:00:39,675
that you can download from 
the resources for this video. 

9
00:00:39,675 --> 00:00:42,720
Lines 2 to 14 give us some drugs to work with. 

10
00:00:42,720 --> 00:00:47,840
I've used tuples, containing the name of the 
drug, and a brief description of what it does. 

11
00:00:47,840 --> 00:00:51,694
Below that – and I'll scroll 
the data up the screen: 

12
00:00:52,400 --> 00:00:55,476
We've got a list containing 
some adverse interactions. 

13
00:00:55,600 --> 00:00:59,760
These are sets, containing drugs that 
shouldn't be taken together. We're not 

14
00:00:59,760 --> 00:01:05,040
going to use this list just yet, but it'll be 
useful when we come to look at subsets, later. 

15
00:01:05,040 --> 00:01:11,056
Okay, I'll scroll through the data 
again, so we can see lines 28 to 39. 

16
00:01:11,760 --> 00:01:16,560
This is a dictionary containing the prescriptions 
for our patients. The keys are the patients' 

17
00:01:16,560 --> 00:01:20,883
names, and the value for each patient is a 
set of the drugs they've been prescribed. 

18
00:01:20,883 --> 00:01:25,360
Please note that we're not medically 
trained, and this is just example data. 

19
00:01:25,360 --> 00:01:29,040
Also note that we're in danger of killing a 
few of these patients – as you'll see if you 

20
00:01:29,040 --> 00:01:34,480
compare their prescriptions with the drugs that 
have adverse interactions. That's intentional. 

21
00:01:34,480 --> 00:01:38,932
Well, killing them isn't intentional.
The mix of drugs is intentional, 

22
00:01:38,932 --> 00:01:43,281
so that we can see how to identify the 
problem prescriptions, in a later video. 

23
00:01:43,520 --> 00:01:49,134
Kenny, in particular, is going to have a hard 
time. Hopefully, if we get our code right, 

24
00:01:49,134 --> 00:01:52,657
he'll survive.
Alright, that's our data. 

25
00:01:52,960 --> 00:01:57,760
Next, we'll use a sample of patients, 
and replace their Warfarin with Edoxaban. 

26
00:01:57,760 --> 00:02:05,376
Create a new Python file called prescription_trial

27
00:02:07,850 --> 00:02:11,089
We'll start with the code to import the data:

28
00:02:16,160 --> 00:02:19,760
We'll look at import in more 
detail, in a later section. 

29
00:02:19,760 --> 00:02:23,840
When we do, you'll see why using 
import * should generally be avoided. 

30
00:02:23,840 --> 00:02:27,840
But it makes our code simpler, for this 
example, and I don't want to confuse 

31
00:02:27,840 --> 00:02:34,585
things with strange dot notation, at this stage.
Next, we'll select some patients for the trial. 

32
00:02:34,585 --> 00:02:36,598
I'll store their names in a list:

33
00:02:52,482 --> 00:03:00,248
Our main code will iterate over this list of patients, and replace their Warfarin with 
Edoxaban. I'll add a comment, then the for loop: 

34
00:03:08,240 --> 00:03:11,556
We get the patients' prescriptions 
from the patients dictionary. 

35
00:03:11,556 --> 00:03:15,990
Remember that the patient's name is the 
key, and their prescription is the value. 

36
00:03:23,360 --> 00:03:28,684
If you're not sure what we've got there, print 
out prescription to check. Let's do that: 

37
00:03:36,720 --> 00:03:39,080
Run the program, to check what we're getting.

38
00:03:40,071 --> 00:03:42,240
Okay, for each patient, we get the 

39
00:03:42,240 --> 00:03:47,280
set of medications that they've been prescribed.
All four patients are taking Warfarin – remember 

40
00:03:47,280 --> 00:03:51,760
that sets aren't ordered, so warfarin might 
not be the last item for each patient. 

41
00:03:51,760 --> 00:03:56,000
To perform the trial, we'll delete 
Warfarin and replace it with Edoxaban. 

42
00:03:56,000 --> 00:03:59,385
First, we remove Warfarin 
from the prescription set: 

43
00:04:05,492 --> 00:04:07,690
Then we add Edoxaban instead:

44
00:04:13,876 --> 00:04:17,291
Run the program and check the output.

45
00:04:17,588 --> 00:04:21,220
Each patient now has Edoxaban instead 
of Warfarin in their prescriptions. 

46
00:04:21,360 --> 00:04:27,333
That's working fine. On line 8, we use the remove 
method to remove warfarin from the prescription. 

47
00:04:27,520 --> 00:04:32,688
On line 9, we add edoxaban.
Now let's see why discard isn't appropriate here. 

48
00:04:32,880 --> 00:04:36,433
I'll replace remove with discard, on line 8:

49
00:04:44,598 --> 00:04:47,243
Run the program, and everything still works.

50
00:04:48,400 --> 00:04:51,600
So what's the problem?
To see what could easily happen, 

51
00:04:51,600 --> 00:04:57,200
let's see what our code will do if Kenny gets 
included in this trial. Remember Kenny? He's one 

52
00:04:57,200 --> 00:05:02,458
of the patients that we're in danger of killing.
I'll add him to the list, on line 3: 

53
00:05:22,560 --> 00:05:26,455
Run the program again, and check 
Kenny's prescription carefully. 

54
00:05:26,800 --> 00:05:31,375
He hadn't been prescribed an anti-coagulant, 
but he's now going to be taking Edoxaban. 

55
00:05:31,680 --> 00:05:36,000
Now, the patients for this trial should have been 
selected at random, from only those patients who 

56
00:05:36,000 --> 00:05:41,440
were prescribed Warfarin. But mistakes happen.
Our job is to make sure we don't make things 

57
00:05:41,440 --> 00:05:45,328
worse, by letting our code prescribe 
a drug that someone shouldn't take. 

58
00:05:45,520 --> 00:05:48,401
This is where remove is more 
appropriate than discard. 

59
00:05:48,560 --> 00:05:53,622
Let's change the code back, to use remove 
on line 8 instead, and see what happens: 

60
00:06:00,880 --> 00:06:06,524
Run the program, and it crashes.
We get a KeyError on line 8.

61
00:06:06,720 --> 00:06:10,657
warfarin wasn't found in Kenny's 
prescription, so it can't be removed. 

62
00:06:10,800 --> 00:06:15,840
Now that's not good! We shouldn't write code 
that might crash. But I think you'll agree that 

63
00:06:15,840 --> 00:06:20,080
a crash is definitely better than killing Kenny.
You may be thinking "why don't we just check 

64
00:06:20,080 --> 00:06:25,791
if warfarin is in the set, before we try to 
remove it?". And of course, we could do that. 

65
00:06:26,000 --> 00:06:29,122
I'll add the condition, before line 8:

66
00:06:33,741 --> 00:06:38,760
and indent lines 9 and 10.

67
00:06:39,840 --> 00:06:42,210
Then I'll add an else block:

68
00:07:07,298 --> 00:07:09,008
That fixes the problem.

69
00:07:09,200 --> 00:07:11,659
When we run the program: 

70
00:07:12,880 --> 00:07:18,227
Kenny is no longer prescribed Edoxaban. We also 
get a message to remove him from the trial. 

71
00:07:18,400 --> 00:07:22,560
However, this isn't an ideal 
solution. Performing that test, 

72
00:07:22,560 --> 00:07:27,523
on line 8, takes time – testing 
conditions is a relatively slow operation. 

73
00:07:27,760 --> 00:07:32,100
Our trial could include thousands, 
or tens of thousands, of patients. 

74
00:07:32,240 --> 00:07:36,160
We're performing that test for every single 
patient, even though it will be unnecessary 

75
00:07:36,160 --> 00:07:42,594
for most of them. Any patient not taking Warfarin, 
who got included in our trial, is an exception. 

76
00:07:42,800 --> 00:07:48,627
There shouldn't be many of them. In fact, there 
shouldn't be any of them, but mistakes can happen. 

77
00:07:48,800 --> 00:07:52,000
So we've slowed down our code, with 
a test that will only be False for 

78
00:07:52,000 --> 00:07:57,360
a very small number of patients, if any.
The more Pythonic way to deal with this situation, 

79
00:07:57,360 --> 00:08:03,760
is to use an exception. Kenny, after all, 
is an exception. He's not taking Warfarin, 

80
00:08:03,760 --> 00:08:07,283
and managed to be included in 
this trial because of some mix-up. 

81
00:08:07,360 --> 00:08:10,500
We'll be covering exceptions, in a later section. 

82
00:08:10,640 --> 00:08:15,107
When you start using them, it's even more 
obvious why remove is more appropriate here. 

83
00:08:15,280 --> 00:08:19,440
Because we haven't covered exceptions yet – 
we can't cover everything all at once – sit 

84
00:08:19,440 --> 00:08:23,627
back and relax, while I demonstrate 
how catching the exception would work. 

85
00:08:23,760 --> 00:08:28,020
Don't worry about understanding this 
code. You'll learn all about it soon. 

86
00:08:28,160 --> 00:08:31,680
But it will show why it can be useful to 
use a method that raises an exception, 

87
00:08:31,680 --> 00:08:35,120
as the remove method does.
I'll change line 8, 

88
00:08:35,120 --> 00:08:38,595
changing the condition to a try instead: 

89
00:08:41,520 --> 00:08:45,816
Next, the else block is replaced 
with an exception handler: 

90
00:08:50,560 --> 00:08:54,400
The code doesn't look very different, 
but we're no longer testing a condition, 

91
00:08:54,400 --> 00:08:58,320
each time round the loop.
For most patients, lines 9 and 10 

92
00:08:58,320 --> 00:09:04,400
will execute fine. For the very few patients (if 
any) who shouldn't be in the trial, the exception 

93
00:09:04,400 --> 00:09:10,247
will cause lines 12 and 13 to execute.
I'll run it, to make sure it works. 

94
00:09:10,560 --> 00:09:13,600
And we get the same output.
This works because the remove 

95
00:09:13,600 --> 00:09:17,888
method raises an exception, when the 
item we try to remove isn't in the set. 

96
00:09:18,000 --> 00:09:22,879
And when you start using exceptions in your 
code, that becomes an important difference. 

97
00:09:23,120 --> 00:09:28,195
I'll finish this video by seeing what 
happens, if we used discard instead of remove: 

98
00:09:35,462 --> 00:09:40,480
I've changed remove to discard, on 
line 8. As we've seen, discard doesn't 

99
00:09:40,480 --> 00:09:45,460
give an error – raise an exception, in 
Python speak – so Kenny's in trouble. 

100
00:09:45,600 --> 00:09:50,566
Run the program ... and sure 
enough, we've killed Kenny. 

101
00:09:50,880 --> 00:09:57,071
As I said, don't worry about that try/except code. 
We'll be learning all about exceptions, later. 

102
00:09:57,200 --> 00:10:00,960
I included it, because it really demonstrates 
why you might want a method to raise an 

103
00:10:00,960 --> 00:10:05,187
exception – an error in other words 
– rather than silently doing nothing. 

104
00:10:05,280 --> 00:10:09,200
I'll change discard back to remove, 
on line 9, because the try block 

105
00:10:09,200 --> 00:10:10,494
doesn't make sense otherwise:

106
00:10:18,510 --> 00:10:20,209
Alright, in the next video,

107
00:10:20,308 --> 00:10:27,280
we'll have a look at another method that deletes 
an item from a set. See you in the next video.

