WEBVTT 1 00:00:01.900 --> 00:00:05.660 so going back looking at our code again if we opened up song.py and I just closed down 2 00:00:05.660 --> 00:00:11.700 the run window incidentally I just want to point out that the line on 132 the 3 00:00:11.700 --> 00:00:16.209 print that format doesn't work with Python 2 so in another words Python 2 doesn't 4 00:00:16.209 --> 00:00:20.750 allow printing in that way but you can import the print function from the 5 00:00:20.750 --> 00:00:24.970 future module if you wanted to work again just for Python 2 so if your running Python 6 00:00:24.970 --> 00:00:29.619 3 you don't need to do this but I'm just going to quickly show how to that fi you happen to be running Python 2 7 00:00:29.619 --> 00:00:32.750 just come up here onto the first line and add an import 8 00:00:34.530 --> 00:00:49.570 .... 9 00:00:49.570 --> 00:00:54.280 intelligence saying are you using something from an older version of Python so click on no 10 00:00:54.280 --> 00:00:55.680 for now 11 00:00:55.680 --> 00:00:58.850 so from __future__ 12 00:00:58.850 --> 00:01:06.550 import....so if you added that to your Python code and added 2 blank 13 00:01:06.550 --> 00:01:10.190 lines for Python 2 you find that this should work and that print statement that 14 00:01:10.190 --> 00:01:16.660 I pointed out should work on Python 2 but I'm gonna delete that now because of course we're dealing with Python 3 15 00:01:16.660 --> 00:01:21.600 so we ran it and then we had a quick look at albums.txt and check.file.txt 16 00:01:21.600 --> 00:01:27.550 in the previous video and they seem to be the same so what we can do as I alluded to 17 00:01:27.550 --> 00:01:30.150 we can compare 2 files in IntelliJ 18 00:01:30.150 --> 00:01:34.720 so the way to compare two files you just select them so albums.txt and holding 19 00:01:34.720 --> 00:01:39.409 command and clicking check file and holding control if you're on Windows or 20 00:01:39.409 --> 00:01:44.310 Mac and once you've selected both here on the left hand side you can come up here to the view 21 00:01:44.310 --> 00:01:51.510 menu and click on compare files or could have done a command D or a control d on Windows or Linux so 22 00:01:51.510 --> 00:01:58.450 click on compare files and its open up a window and has put contents are identical you can see 23 00:01:58.450 --> 00:02:01.240 that's actually compared with 24 00:02:01.240 --> 00:02:05.080 or compared both files and incidentally just going back to the code and if 25 00:02:05.080 --> 00:02:08.619 you select one file click one like that 26 00:02:08.619 --> 00:02:14.060 you can then go to view and click on compare with the second file you can actually grab 27 00:02:14.060 --> 00:02:17.590 that file literally from another project if you wanted to compare it that 28 00:02:17.590 --> 00:02:21.720 way so that's really useful if your comparing files that are in different projects such as your own 29 00:02:21.720 --> 00:02:25.640 code compared to the code downloaded from the resources for this course for 30 00:02:25.640 --> 00:02:30.250 example so you don't have to have the 2 files comparing within the same file but just having 31 00:02:30.250 --> 00:02:34.849 a look at that the compared results their again looking at this it's probably a 32 00:02:34.849 --> 00:02:38.400 good indication that the program is reading the data correctly but it 33 00:02:38.400 --> 00:02:43.160 doesn't prove at the moment and the program doesn't have any errors and as 34 00:02:43.160 --> 00:02:47.459 it turns out there's actually quite a serious flaw in the program but the test 35 00:02:47.459 --> 00:02:51.629 data doesn't reveal it and it's very important to test program with a range 36 00:02:51.629 --> 00:02:56.450 of different inputs and be very suspicious of data that is to neat such as the 37 00:02:56.450 --> 00:03:05.609 example take albums.txt file so we just close this comparison down and go back to the albums.txt 38 00:03:05.609 --> 00:03:09.720 examining that files it's nicely ordered with the album for each artist following on 39 00:03:09.720 --> 00:03:13.859 from each other and all the songs for an album group together but real world data 40 00:03:13.859 --> 00:03:18.500 often and probably usually isn't like that and so we should at least probably scramble the 41 00:03:18.500 --> 00:03:22.450 file up a bit to be able to test this properly now it doesn't take much 42 00:03:22.450 --> 00:03:25.950 scrambling to reveal the errors so we are gonna edit albums.txt 43 00:03:25.950 --> 00:03:30.950 and what i'm gonna do is cut out all the songs for zz top albums so come down here to 44 00:03:30.950 --> 00:03:43.049 zz top right down the bottom so what I'm gonna do is I'm going to cut out all the songs zz top album antenna 45 00:03:43.049 --> 00:03:51.410 this line 657 down through 668 so cut those command x on a Mac and 46 00:03:51.410 --> 00:03:58.690 make sure I delete the extra line as well be control X on Windows or PC and what I'm gonna do 47 00:03:58.690 --> 00:04:03.790 is I'm going to actually put them in a different spot in the file so go back up back 48 00:04:03.790 --> 00:04:08.989 up to the top and put them somewhere completely different say aerosmith will do so 49 00:04:08.989 --> 00:04:10.040 let's put it 50 00:04:10.040 --> 00:04:16.699 somewhere around here so maybe between permanent vacation and jim so just 51 00:04:16.699 --> 00:04:23.020 here lets just put it in their so we are putting it in a different order now just to see what happens if we try and run this 52 00:04:23.020 --> 00:04:26.789 and again I made sure that there's no blank lines either the lines I've 53 00:04:26.789 --> 00:04:32.710 added here or back where I cut it out which was on lines 650 making sure 54 00:04:32.710 --> 00:04:37.100 that there is no other lines there and there is a blank line on the line at 55 00:04:37.100 --> 00:04:41.590 the end of line 685 which I have left I haven't touched that so we've now split 56 00:04:41.590 --> 00:04:46.660 two artists albums up a little bit remembering that last time when we ran this 57 00:04:46.660 --> 00:04:52.210 it resulted in twenty-eight artists so lets run this again and we got a different 58 00:04:52.210 --> 00:04:55.550 result as you can see here we've now got their are 30 artist and again 59 00:04:55.550 --> 00:05:00.530 remembering those 28 last time when we ran this in the previous video and 60 00:05:00.530 --> 00:05:03.729 we just go and do that comparison again 61 00:05:03.729 --> 00:05:08.849 so albums.txt and check file selecting them both again and click on view and compare 62 00:05:08.849 --> 00:05:14.530 files and just open that compare files results but it still says in the top left hand corner 63 00:05:14.530 --> 00:05:18.460 contents are identical but its now reporting 64 00:05:19.010 --> 00:05:23.010 their is 30 artists so obviously somethings is badly wrong here because we only had 65 00:05:23.010 --> 00:05:27.419 twenty eight artists in the previously video we've now got 30 but yet the output is still the 66 00:05:27.419 --> 00:05:32.710 same so the point that we're making here is really does illustrate how important 67 00:05:32.710 --> 00:05:38.820 it is when you testing to test with varied data sets ok so that is the problem so onto 68 00:05:38.820 --> 00:05:41.800 what or how to go about fixing it 69 00:05:41.800 --> 00:05:47.210 problem is that each artists boundary in the data causes a new artist object to be 70 00:05:47.210 --> 00:05:51.690 created without considering whether there's already an object for that artists 71 00:05:51.690 --> 00:05:55.710 so as a result we end up with two objects for Aerosmith and to for zz top 72 00:05:55.710 --> 00:05:59.139 as it turns out the code to fix this is quite easy 73 00:05:59.860 --> 00:06:03.510 you really just have to check to see it is already an artist in the list before 74 00:06:03.510 --> 00:06:07.780 creating a new artist object and to do the same each time a new albums found 75 00:06:08.400 --> 00:06:12.110 but rather than fixing this version though we're going to go ahead and use a 76 00:06:12.110 --> 00:06:13.699 slightly different approach 77 00:06:13.699 --> 00:06:16.880 now the way that this program currently works does have some 78 00:06:16.880 --> 00:06:20.710 disadvantages including the need to check if there's an outstanding artist 79 00:06:20.710 --> 00:06:28.229 or album record at the end of the this is the code down here on line 116 onwards 80 00:06:28.229 --> 00:06:32.320 so we have to do that final check for the last line so the method uses 81 00:06:32.320 --> 00:06:36.690 easy-to-understand and sometimes it is necessary to do things this way as an 82 00:06:36.690 --> 00:06:39.770 example with restoring the data in a data base and we might want to ensure that 83 00:06:39.770 --> 00:06:45.150 an artist had at least one album before adding them to the database now with some data such as 84 00:06:45.150 --> 00:06:49.830 financial records it may also be a requirement to wrap the database 85 00:06:49.830 --> 00:06:54.810 insertions into a transaction so that either all the data of an artists including 86 00:06:54.810 --> 00:06:59.000 albums is written or not if it is so that way there's no chance of incomplete data 87 00:06:59.000 --> 00:07:03.150 being stored in the database similarly you may need to make sure that 88 00:07:03.150 --> 00:07:07.270 all the song for an album are added when the album is added or abandoned the attempts 89 00:07:07.270 --> 00:07:12.229 to store the data altogether so taking this approach of not adding the object to their lists 90 00:07:12.229 --> 00:07:15.979 until the related data has been read from a file is one way to sort of 91 00:07:15.979 --> 00:07:20.160 achieve this that doesn't really apply to this particular example though because 92 00:07:20.160 --> 00:07:23.840 all the data will fit in memory but with a data file that was to big for available 93 00:07:23.840 --> 00:07:27.900 memory would really be necessary to store the records as soon as everything 94 00:07:27.900 --> 00:07:32.159 related to an object had been read so although the algorithm that we've used may 95 00:07:32.159 --> 00:07:36.270 not be the best for this particular case it can be useful in other cases and that's 96 00:07:36.270 --> 00:07:42.080 why that we've actually included it here and algorithm is a series of steps 97 00:07:42.080 --> 00:07:45.490 for solving problems so I've described the algorithm for this approach briefly 98 00:07:45.490 --> 00:07:50.740 before writing the load_data function but I'm gonna bring up a flow chart to 99 00:07:50.740 --> 00:07:56.010 show this little bit more formerly and I'll just bring it up on the screen their now you can see I've got flow chart here 100 00:07:56.010 --> 00:08:01.370 so this is really describing what's what's really happening now with the 101 00:08:01.370 --> 00:08:05.370 flow of the program code and it can be useful if you haven't seen these before 102 00:08:05.370 --> 00:08:09.990 to investigate and find out more about flow charting when your trying to 103 00:08:09.990 --> 00:08:14.820 put together how something works it can be useful to put this together but you can see what 104 00:08:14.820 --> 00:08:19.279 is happening here is that we're starting the process it's reading some data it is 105 00:08:19.279 --> 00:08:21.460 then splitting the data into the various fields 106 00:08:21.460 --> 00:08:26.930 and you've seen that before and we start with a comparison here is the artist the same as the field 107 00:08:26.930 --> 00:08:32.830 field being what we have read in if it is the same then it flows over to here we do another check is the album the same as the 108 00:08:32.830 --> 00:08:33.890 field 109 00:08:33.890 --> 00:08:38.680 if it is we are going to add song to album checking to see if there's any more data if their is it goes back and 110 00:08:38.680 --> 00:08:42.570 sort of reloops through and in the case of their wasn't anymore data will save the 111 00:08:42.570 --> 00:08:46.430 artists and then we stop bullets so that is in that scenario their and that is with the 112 00:08:46.430 --> 00:08:51.110 artists was the same as the field but if the artist isn't the same as the field then we 113 00:08:51.110 --> 00:08:57.140 save the artists and we created that artist object and we move to here and we go through this decision 114 00:08:57.140 --> 00:09:00.920 again to see to check whether the album's the same as the field it is 115 00:09:00.920 --> 00:09:04.339 it comes over to hear then adds the song to the album object and sort of flows 116 00:09:04.339 --> 00:09:07.750 though their that we have talked about before that sort of process to be reading more 117 00:09:07.750 --> 00:09:12.250 lines of data or we reach the end to save the artist and stop but in the case of the 118 00:09:12.250 --> 00:09:16.160 album not being the same as the field we save the album and the artist 119 00:09:16.160 --> 00:09:21.860 object we created the album object then come up back up to here and then we 120 00:09:21.860 --> 00:09:25.730 add the song to the album object and go through to this point again so basically 121 00:09:25.730 --> 00:09:30.250 artists objects isn't saved until the new artists is found in the data set or 122 00:09:30.250 --> 00:09:34.390 when their no more data to read now as I mention this could be a reasonable approach to 123 00:09:34.390 --> 00:09:39.230 take in some circumstances but I'm gonna change things so that new objects are 124 00:09:39.230 --> 00:09:44.100 save as soon as they created so we are not using a database here so by save what I mean is 125 00:09:44.100 --> 00:09:48.649 their stored in their respective lists so the flowchart of that is going to be a bit 126 00:09:48.649 --> 00:09:56.050 different so let me bring them up on the screen so close that one down here is the 2nd 127 00:09:56.050 --> 00:09:59.839 flow chart so the main difference with this one and I'm not gonna go through it all then 128 00:09:59.839 --> 00:10:03.800 now but you can sort of go through that at your leisure but outlining the major difference and 129 00:10:03.800 --> 00:10:08.620 that's that the objects are save as soon as they created and that means that an 130 00:10:08.620 --> 00:10:13.290 artist we put into the list before any albums are added to it likewise an album 131 00:10:13.290 --> 00:10:17.070 object is going to be added to the artists list of albums before any songs have been 132 00:10:17.070 --> 00:10:21.930 added to it now if we were using a database to store these details what 133 00:10:21.930 --> 00:10:26.079 that means is that if something goes wrong we could end up with albums in a 134 00:10:26.079 --> 00:10:30.510 database for with some songs missing but for this example though you know when where 135 00:10:30.510 --> 00:10:33.459 we're reading everything into memory the code using the new 136 00:10:33.459 --> 00:10:37.470 algorithm ends up being a lot simpler so I'm gonna add some extra steps to check 137 00:10:37.470 --> 00:10:42.179 if an artist or album already exists and if they do then the existing objects can 138 00:10:42.179 --> 00:10:47.009 be retrieved from the list instead of creating a new one so to do that 139 00:10:47.009 --> 00:10:51.929 lets introduce a slightly more complicated version so lets bring up the flow chart for that so I'll 140 00:10:51.929 --> 00:10:57.709 close this one down and lets open up the third flow chart and you can see the new 141 00:10:57.709 --> 00:11:01.610 sequence here to check to see whether the artists already been stored if it has 142 00:11:01.610 --> 00:11:05.350 only been stored its gonna retrieved the artists and flow on and 143 00:11:05.350 --> 00:11:09.339 likewise when we tested for the album is the album already stored if yes we are gonna retrieve the 144 00:11:09.339 --> 00:11:12.420 album's first and then move on to that if it hasn't of course we are going to 145 00:11:12.420 --> 00:11:18.139 create the album object and of course over here if the artist hadn't been stored then we create it 146 00:11:18.139 --> 00:11:22.439 which we need for the first time one thing that probably stands out looking 147 00:11:22.439 --> 00:11:26.509 at this flowchart is that the new steps that have been added into the flow for 148 00:11:26.509 --> 00:11:31.220 both artist and album practically identical and this sounds good candidate 149 00:11:31.220 --> 00:11:36.189 for function so I'm gonna start by and going back to the code and looking at adding a 150 00:11:36.189 --> 00:11:41.100 find object function before load data but we will start working on that in the next 151 00:11:41.100 --> 00:11:41.360 video