State of Wagtail
Published July 18, 2024
This video features Tom Dyson at DjangoCon US 2018 in San Diego, California, USA.
DjangoCon US 2018 - Here Come The Robots - Django and Machine Learning by Tom Dyson
Machine Learning is probably the most important development in our industry (and possibly our civilisation!). Previously restricted to math geniuses with access to supercomputers and massive data centres, machine learning tools are increasingly available as web services which are easily consumed from more traditional web applications. Python has become the lingua franca of machine learning, so Django developers are well placed to take advantage of the next wave of application development.
In this talk I’ll outline the various machine learning platforms and provide a set of practical examples that demonstrate how Django developers can start taking advantage of artificial intelligence in their own applications. These will include:
Image recognition - using Microsoft Azure Vision to automatically caption and label the images your users upload
Entity analysis - using the Google Cloud Natural Language API to tag news articles with people, locations and events
Predictions - using Amazon Machine Learning to build a ‘you may also like’ feature
Sentiment analysis - using IBM Watson to understand the tone of comments submitted to your site
This talk was presented at: https://2018.djangocon.us/talk/here-come-the-robots-django-and-machine/
LINKS:
Follow Tom Dyson 👇
On Twitter: https://twitter.com/tomd
Official homepage: https://torchbox.com
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Tom Dyson distinguishes machine learning from the broader field of artificial intelligence, then explains machine learning as using data and known answers to derive rules. He demonstrates accessible cloud services for image recognition, sentiment analysis, entity extraction, and predictive classification, showing how Django and Wagtail projects could use them for accessibility, content management, customer support, recommendations, and audience segmentation. He argues that these tools are already practical for developers, but warns that models learn biases and accidental correlations from their training data, making careful data selection and human responsibility essential.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: So uh yeah I'm Tom. I'm uh I work at Torchbox in the UK. We uh we make the Wagtail CMS, we we created Wagtail CMS that uh Thank you. I hope some of you are familiar with and clearly you are. And uh usually that's what I talk about, but um this time somebody else did a much better job. I hope some of you went to Sarah Hines ' amazing talk yesterday And today I'm going to talk about machine learning, which uh which I think is I think it's the most important topic in our industry. Um I want to start with a request and an apology. The request is The request is that some of you find an image that you can share with me on Slack
Speaker 1: when the time comes, and this is going to be part of a live demo. All I ask is the image is publicly accessible and that it's uh safe for work. And uh and the apology is that um the apologies for my clickbaity title. And uh this this is this is not a talk about um kind of Terminator style robots coming in and and taking over and Django fighting them or something. This is a uh it's a bit more prosaic, but um And I think uh the but I I brought I I included it because I feel like there is a confusion between these concepts. And this is something that came out when I was uh describing this talk to some of my colleagues. And I think it's an understandable confusion because artificial intelligence in particular is a term that tends to get really overloaded in the media.
Speaker 1: But um For the most part, this this is not a talk about artificial intelligence. Artificial intelligence is uh I mean the simple thing if you if you want to remember this is artificial artificial intelligence is like a superset of technologies and concepts which includes machine learning. Machine learning is a very important component of it, but it also includes things like natural language processing. So uh highlight for me for the conference so far was the talk yesterday about gender bias and Harry Potter. Anyone come to that? And um you know that and so that that used some really interesting natural language processing techniques to to identify A kind of range of connections between words and sentiments. And machine translation is another example, robotics. And there are many of these technologies. The one that we're going to focus on is machine learning in this talk. But if you're interested in then the wider questions around artificial intelligence, I recommend this book, Life 3.
Speaker 1: 0, by Max Tegmark. He's a a physics professor from MIT, but he writes in a very accessible and compelling way. And uh his thesis is that life as we know it is gonna be um is gonna come to be described in in three phases. So version 1. 0 is uh this is like the the the original, the simplest form of life. This is um this could be as like a single-cell amoeba or a uh a chicken And um and what's common about about these things is that uh they evolve slowly. They evolve through evolution. So their hardware, their bodies evolve. But also the software, the way that they navigate the world, the way that they they make decisions, just happens, you know, very slowly, generation by generation.
Speaker 1: And then the next big change was was us humans. And like chickens, we have relatively little control over our bodies. So uh we can go to the gym or we can eat a lot and we can make sort of temporary changes. But we can't become a hundred times stronger or faster. And again, we can we we're we're we're restricted by by the speed of evolution. However, what's special about humans is that we've been able to design our own software so we can decide uh what language to speak, or we can decide What kind of deep sphere of knowledge that we're going to we're going to become expert in? We can decide where in the world we want to live. You know, some people are at least are able to decide that because of their circumstances.
Speaker 1: And this is a very dramatic change Interestingly, there's a sort of 2. 1, just to slightly muddy the waters, which is that uh in the last 50 or so years, we have been able to make minor changes to our own hardware. We've been able to replace our knees or uh or I could get get new teeth. But it's it's kind of chipping away at the edges. It's not dramatic changes. And then the big one, the subject of the book is Life 3. 0. So 3. 0, like us, can um uh can design its own software, but uh the incrucial difference here is that it can also design its own hardware. So version three of Life will also be able to make A processor that's uh okay a hundred times faster. Um
Speaker 1: are we gonna try resetting this?
Speaker 2: Yeah, that's just Let's keep it.
Speaker 1: Cool, thank you. And when this happens, then we might start seeing some very, very kind of accelerated pace of development. Uh unfortunately life 3. 0 is not us. Life 3. 0 is is the machines. And um to get to life 3. 0 a lot of very hard stuff is going to happen. I mean this is this is going to be it's an extraordinarily hard. problem to solve and uh it's often this is described in terms of AGI, which is uh general artificial intelligence. So Computers are really good. Computers are better at us now at playing chess or driving cars or calculating prime numbers But what they lack is our extraordinary versatility. So you know in this room
Speaker 1: we can probably speak ten languages, we can probably play twenty musical instruments, we probably have like deep knowledge in the fundamentals of the Django ORM or uh you know how to make the best fish taco. And you know, th this this kind of range of skills we have is extraordinary. And actually On my basic understanding of this so far, it feels like there's gonna have to be it's not even really clearly known yet how we're gonna move from the sort of narrow intelligence that that uh that's that's already working really well to this general intelligence. Nevertheless , okay, it's not true to say there's a consensus about where this will happen, but in the in the last survey of experts in this field, the median answer uh at the point at which we achieve AGI, general intelligence, is 2050.
Speaker 1: So there are a lot of people saying it's going to take a lot longer, some people say it's going to take sooner. The media answer is is is 2050. And at that point, particularly when you get to human human level AGI, sometimes AGI just means human level. Then something really interesting happens because from that point you only need to get to kind of 1% more than human and then then things may just kind of become recursive and rapidly increase because a machine that's one percent smarter than us is going to be better than us at building machine learning computers. And then, you know, and so on and so on. So that's that's the sort of the the trigger point. That's at least that's the idea. I think it's uh it feels like kind of an amazing time to be alive, right?
Speaker 1: Life 1. 0, as far as we know, started around 4 billion years ago. And then humans have been around. We've been around for about uh 100 millennia. And then this next phase is going to happen in 30 years, possibly. You know, that's that's the median answer. In most of our lifetimes, we're going to see this this massive change. Um so I don't know, it's uh f it's it's it it's both frightening and and and exciting. It feels to me like particularly striking that the same timescales that we're talking about this life 3. 0 are the timescales in which most scientists agree that we may enter the kind of cataclysmic results of climate change. And I don't know if that's just like a sort of cosmic joke that the
Speaker 1: kind of the uh the you know the solution may arrive just at the same time or or just after. Or or maybe it's just like a kind of replaying of the same uh kind of Apocalypse versus Messiah story that's been prevalent in world religions for uh for millennia too. Anyway, that's the end of the kind of the uh metaph metaphysical bit of my talk, and the rest is uh is a bit more about Relevance to us as uh Django developers and focusing particularly on machine learning, this kind of specific area of of artificial intelligence. I really like this definition of machine learning by Francois Cholet And the classical programming, so the programming that we do day in, day out, uses rules and data to produce answers, whereas machine learning uses data and answers to produce rules.
Speaker 1: I think a good example of this is uh with spam. So how many of you uh how many of you used email before Gmail? Alright, so quite a quite a number, more than I was expecting. But so uh uh I did and um and Uh for those of us who who used to have to run their own email servers, so you know if you were worked for a were you know in your organization you would you would have an email server and uh it it You have to deal with spam. So everyone has to deal with this this this problem of spam. And there are tools like Spam Assassin, and you basically start writing these rules. So everyone in my contacts list, that's not spam. Something that you know includes like get rich quick, that's that's that's spam. But then um you know of course the the people write writing
Speaker 1: spam messages are kind of aware of these rules and they start adapting and then you have new contacts and they get marked as spam and you have to you become in this becomes in this arms race of uh you know your rules versus the reality. Um and that's the kind of classical programming technique. And it's, you know, the which which is very effective in many cases, but it but it's hard to deal with that that could those kind of attacks. The machine learning technique is basically so what Google's trick was, and when you know you started using Gmail, it felt like that problem had just gone away. They took a billion messages. and they took, you know, a a load of willing volunteers in the shape of users like like us to to mark spam or not spam. So they took the data and they took the answers and then they created the rules and then they just apply that machine learning model. to say whether this message is spam or not spam.
Speaker 1: So that's a kind of simple definition of these two two approaches. Francois Cholet who who who came up with this really nice pithy quote, he's the author of uh Keras, which is the Python API for uh machine learning. Uh he works at Google. He also wrote this book, Deep Learning with Python, which is a fantastic introduction to these tech this technology. He's also amazing on Twitter and every day he says something Which I feel he's probably the cleverest person alive. Okay, so I'm gonna talk about four different uh four different techniques for machine learning. And the first one is, what is this a picture of? This is something that I think at the first DjangoCon ten years ago, this would have felt like an outlandish request to make of someone. You know, taking this array of pixels with different colours and brightnesses, tell me what the subjects in this pic
Speaker 1: this this are. I mean imagine trying to write rules to to work that out. But now we're I'm able to do this with uh twelve pretty crufty lines of Python. Um I'm not going to kind of describe all of this, but most of this is handling the authentication, sending the right headers, and then pinging a service. In this case I've used Microsoft's uh m uh vision service. Amazon have one called recognition, Google have one, Microsoft have one. You know, uh most of the big kind of cloud providers have these image recognition servers. tag so they'll say human dog. Microsoft will also do a provide description. So I'm just I'm sending the request, I'm getting it back, I'm doing some
Speaker 1: ugly munging because uh I want to just get the captions that are over a particular confidence limit and pick the top one of those. Okay, this is where I'm gonna move to the demo. Right, and this is where I would like uh anyone to send me, anyone who's prepared this, to set to send me uh an image on Slack and we're gonna we're gonna try this out. So it just has to be publicly accessible. Uh his one Oh here's one from Amy. Okay. So this is pretty cool. Do it. Alright. Copy link.
Speaker 1: Describe. There's the picture. A bird sitting on top of it. Alright. Well, I'll scroll down. Um I I don't think it's a bird. It looks like a sort of, you know, it looks a it looks like a baby. Is what it's a sloth. That's a pretty cute sloth. But it is sitting on top of the table, kind of more in the cup of coffee than next to it. Alright, let's have another one. Scott. There's a massive URL. Copy link. Okay, what's this picture gonna be? So waterfall Dot. Dot. That doesn't look too hard. Shh.
Speaker 1: Scott, that's not a dot. Uh sure, no, that's uh no, that's not gonna work. This is mean. Alright, Tim, what's Tim Tim Tim's got? Yeah, copy that link. Okay, it's a nice picture. Welcome to Bummy Beach. Also dot okay this is not this is not a good live demo. Apart from the sloth. Can I try one more? Has anyone got one more?
Speaker 1: Okay. Trouble is that's not an image. That's a link to a context. I'd have to get the image out of that, all right. It's probably not gonna make me copy the copy image address. Let's try that. Well, okay, so you know I think uh Microsoft's gonna be work to do on this one Thank you for all your challenging demos. I did have some pre-prepared ones that look really good, but I thought I'd you know risk it a bit.
Speaker 1: I'm gonna move on. No, I'm not actually. I'm gonna show you a pre-prepared demo. So um so here's here's here's something uh a bit more in context. Can I zoom in on this? Yeah. Alright, this is this Uh Wagtail CMS. So this is some someone who um Swedish guy who works at a cool agency called Freud, I think. Freud. He built this as a plug-in to Wagtail. And in fact, it uses the same service, the the the Microsoft one. So this is just standard Wagtail. I'm taking some images from my desktop. Select them all and drag them in and this is normally the point at which I would tag them. But now we collect we can see that the the descriptions have been put for us. So the bottle of beer sitting next to a glass of wine. Red plate on a table. So the tags are there as well.
Speaker 1: You can see underneath the herd of table. There's my dog, black and brown dog. And I think this is uh apart from the the description's actually working this time. I think this is a really nice example of uh of of how this of how this could play out because clearly you're not going to be able to rely on this this yet. Um and uh you know the reason that uh that I guess it didn't recognize the sloth is because Uh Microsoft's library of images probably doesn't include many baby sloths yet. But that will improve so so that you know the data will get better But at the same time, we can use these techniques, I think, just kind of augment an editorial experience. So, you know, for a situation where a content manager is dealing with a lot of images, then this is something that could speed up that process. What other use cases? So accessibility, I mean providing alt
Speaker 1: tags, that's like, you know, that's the reason that we're doing this and and and websites. you know have a good technique for that HTML has an answer for that as long as you provide the alt tag but there are other situations in which case in which uh you m this may not be so straightforward. So for example you may be sharing images on a Slack channel and you may have colleagues at work who aren't able to see those. You could write a bot that would would take those images and prepare descriptions of them, hopefully more accurately than the ones we've seen before. Most of these most of these uh tools have systems for for telling you who can can give you kind of whether or not this is uh pornographic or not safe for work material. I thought a nice idea might be for language learning. You could build a really simple app that would help people like take a picture and then describe it. And you know, this could be something that you could do as a sort of simple language learning technique. Moving on, because I'm I've only got 10 minutes left, onto sentiment analysis.
Speaker 1: So this is this is uh perhaps a a more common uh a more common kind of uh idea and this is uh trying to understand what the author is feeling when they write this. Um in in this case again it's this like three lines of Python. So I just have to do the authentication bit. I send my uh I send my request and then I and I get the document tone. In this case I'm using IBM's Watson service. IBM are really hot on this because they they're kind of they they want to do a lot around um bots, I think Okay, demo. Destroy the demo. Moving on to sentiment. So a classic use case for uh for sentiment analysis is with reviews. I'm quite I like Mexican food. I want to see what the best Mexican restaurants are in San Diego. The top one on TripAdvisor is this one called the Taco
Speaker 1: Stand. Does anyone know it? Is it good? Yeah. Um so I can see some reviews. Uh if I start by picking uh pick an excellent review, the best burrito of my life. Okay, so that's sounds pretty positive Let's plug that in. What's it gonna say? Tone joy. That's 0. 71. So it's like 80% joy is the sentiment expressed in this. I mean that sounds about right for a good burrito Um okay yeah but I think generally bad reviews are a bit funnier. Um so let's uh we remove the excellent one. Most of them are excellent. Probably is a really good place. Terrible.
Speaker 1: The first one's good, no, but this one's funnier. I really like this one. Scan my credit card. What's nice about this one is the mix of outrage and honesty. Scan my credit card. Tried to add a $40 tip to my $24 bill. Nice try, you thieves. Food was amazing. We'll never go back. Uh yeah, I really I I like that guy. He's really Prepared to uh to be straight about it. So here. So there is joy, but it's tentative. I think that that's that's kind of accurate. Um So what are the use cases? Customer support is the obvious one, and this is the one that of most often comes up. So you know you're getting feedback from customers. You want to know how they're feeling. But I think there's some more interesting ideas like
Speaker 1: news analysis. So you know, if you're interested in something, if you want to like m know whether Bitcoin's going up or down, you might scrape a hundred news articles every day and and measure the sentiment and you know that might give you some indications. Or better bots, so you know bots writing a good bot is all about being realistic. But if you can measure the tone of the person writing to you before you respond then you're more likely to be able to give a real realistic answer. Okay, racing through to number three, entity extraction. This one's a bit more prosaic, but this is about fundamentally what proper nouns is this text about. And you might think you could do this yourself quite quickly using rules like Something letters that start with capital words that start with capital letters are probably proper nouns, but then you have to think what about the ones at the start of the sentences and uh it turns out that this is this is a slightly harder problem than than you'd expect. Um again
Speaker 1: like 20 lines of Python. This one I'm using the Google Natural Language Service. And let's try this one out. Uh so is uh if I get a new uh it's I'm gonna take this text from uh from the conference website. Add this into text and extract. And here are the entities that it's found out. So it it realizes that this this piece is mainly around DjangoCon US and that it's to do with community and it's also to do with Django the thing. Person is a bit less useful. useful. So this is you know a pretty powerful and very accessible piece of content. I've got a really nice another demo demo using this in Wagtail but I'm going to skip this for now Some examples as so as well as content management.
Speaker 1: And in content management, I guess the main use case is like finding thematic links between pieces of content. So helping people, especially when you have like if you have a blog with a hundred articles, then you probably know what the relationships are. But if you're running a new site with uh two million pieces of content, then you're unless less likely to know that this new article is related to these other articles in different ways. Using tools like this helps you create those themes. Plagiarism is a big big issue now in in higher education in particular. You could look at patterns. Okay, finally, this is the kind of perhaps the most interesting one and it's the one that we have most control over. And broadly, this this type of machine learning project is about working out what's going to happen next given what we already know about the past. And for this step, it's not just like you can't just send your bit of content and ten
Speaker 1: lines of Python. You have to do a bit more work. So you prepare your your content, you have to train, then you evaluate how well the model's working, and then you use it For this example I'm using, this is quite a well-known data set, it's a hundred animals, a zoological database, and uh here's a selection of them. We've got the names and then then the attributes, which in machine learning are described as features. So does it have hair? Does it have feathers? Does it have eggs? And this is kind of binary. So we feed all this in. Most of the challenge in doing this, this is I'm and for this one I'm using the Amazon Machine Learning Service. Most of the challenge for this is like getting through the Amazon docs. And uh while I was uh while I was writing this I found this brilliant tweet from Vicky Boykas who uh And and I feel like for in in in what standing up here now, I'm I
Speaker 1: delivered both three and one from her hottest programming skills of 2018. Um but getting info from the AWS documentation was a big challenge. So I I I upload my my CSV, it has to go through S3. And then you have to explain what these fields mean. And on the most hand, Amazon got it right, so it works out that most of these are binary. Made a mistake on class type, that's the final one. This is the answer we want to get. So out of all these features we want to say what kind of animal is it And this should be categorical because it's not it's not like a three is higher than two. Okay, let's try the demo. And this is the last one. So let's give me an uh gimme an animal. Okay. What was that? Okay, okay, you've got to answer then.
Speaker 1: Does an odd vault have hair? No. Feathers? No. Eggs? It has hair. It has hair. It has hair, okay. It has hair. Eggs? Milk? Yes. Airborne? No. Aquatic? No. Predator? No. Yes. Toothed. Backbone? Yes. Breathes? Venomous? Fins? Tail? Yes. Domestic? Is it cat size? Is it a mammal? Well, there you go. All right. So that was that was uh so having created this model and trained it on that data and tested its evaluation, I can then the Amazon then provide an endpoint and I can fire off more questions like that and it will give me
Speaker 1: it will give me the result There's many draw many applications of this, right? This is this is kind of the main field of machine learning. So audience segmentation is the kind of classic commercial one, working out what kind of users you have. So most of our clients are in the nonprofit space. So I think it's a more interesting problem for us is thinking about ways we can do donations better. So you know, are they a Mac user? Where are they based in? Where are they based? subites that they go to beforehand. We can use tools like that to work out whether we're going to ask for a hundred dollars or fifty dollars. I think a nice one is like this last one. You can you don't have to use this on kind of big public data. You can create your own rules based on your own happiness. So you can think What did I eat? How much sleep did I get? What did I work on today? How happier was I at the end of it? You just feed in that a hundred times and then
Speaker 1: you can start learning the lessons of how to be happier. I just want to point out though, and I'm I'm aware I'm out of time, so I'm gonna try and race finish through this quickly. This is a point at which uh for this kind of machine learning there's more responsibility on you to get this right, okay For example, here, this is a very simple model, but you can see that most of the most of the creatures that begin with B are category four, mammals If you just gave that data to Amazon, it might think that if it starts with B, it's more likely to be a mammal. We know that's correlation, not causation, but you have to, you're responsible for telling it that And there are some much more sinister and dangerous areas in which machine learning can can can get this wrong. This is a story that came out last week. Amazon had this AI recruiting tool.
Speaker 1: This is an extraordinary story. This is a quote from someone who was in involved at Amazon at the time. They wanted it they wanted an engine where you give it a hundred resumes and it will spit out the top five and we'll hire them. But they discovered that it was basically it was not rating them in a gender neutral way and it's because it was trained to vet applicants by observing patterns Submitted to the company from over a hundred from over a ten-year period. Basically, you know, and then and then you get these headlines that AI is evil. AI, you know, AI is not evil. Humans have been evil and prejudiced, and if you train computers on that historical data in which we're exhibiting those mistakes, then computers are going to continue continue making them. So we have to be really careful and and there are some responsibilities on us to get this right. This is another this is a famous post by uh uh Robin Speer, uh where she did the kind of the minimum
Speaker 1: um built the minimum tools necessary to use like uh a standardly available corpus of text from from news stories and then measured sentiment across it and then you know just these examples here show that the correlations between nationalities and positivity has this you know really strong bias and it's you know it's it's it's very easy for us to to make mistakes by relying on past data and perpetuating prejudices. Okay, finally, finally, if you want to learn machine learning, I really recommend that book, Deep Deep Learning with Python and Kaggle's fantastic community. If you want to do something straight away with machine learning, you don't need to do that. You can just read the docs of those cloud services. It took me three hours to put those sites together. And I hope you all build something amazing.
Speaker 1: With it, thank you.
Traditional programming applies rules to data to produce answers. Machine learning uses data and known answers to derive the rules itself.
Discussed at 8:20Cloud vision services can identify objects, generate descriptions, and add tags to uploaded images. In Wagtail, this can speed up editorial work and help provide alt text, while still requiring human review.
Discussed at 15:21Sentiment analysis estimates the tone or emotion in text, such as the joy or anger in a restaurant review. Possible applications include customer support, analyzing news trends, and helping bots respond more naturally.
Discussed at 19:59Entity extraction identifies the proper names and subjects discussed in a piece of text. It can help connect related articles thematically, especially across very large archives, and can also be used to detect plagiarism patterns.
Discussed at 20:30You prepare the data and its features, train a model, evaluate how well it performs, and then use the resulting endpoint to make predictions. Tom demonstrates this by predicting an animal category from attributes such as hair, feathers, eggs, and backbone.
Discussed at 22:18A model can learn correlations and prejudices present in its historical training data rather than genuine causes. The speaker cites Amazon’s recruiting tool, which reproduced gender bias because it learned from past hiring patterns.
Discussed at 25:19You can begin by following the documentation for cloud machine-learning services; Tom says his example sites took about three hours to build. For deeper study, he recommends Deep Learning with Python and the Kaggle community.
Discussed at 26:51Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026