Full-Text search in Django: a journey between different approaches

This video features Ashley Grealish and Marco Alabruzzo at Django London 2020 in Online.

Full-Text search in Django: a journey between different approaches
0:31:41
Published November 27, 2021
183 views

(Meetup starts at 05:32, talk starts at 14:08)

Have you ever wondered how Python's dictionaries work behind the scenes? For the curious minds: we will unveil some of the magic, things ranging from performance to security, and some surprises. For the pragmatists: we'll see cases where understanding the internals can have practical applications.

Presented August 2020 at the London Django Meetup: https://www.meetup.com/djangolondon/events/272166227/

Subtitles kindly provided by Ashley Grealish

Summary

Full-text search is useful because categorising data manually is slow and biased, while search lets users find information directly. Marco compares Django’s standard text lookups, external engines such as Elasticsearch, Solr and Algolia, PostgreSQL full-text search, and PostgreSQL trigrams. Simple lookups suit internal admin filtering; external services provide scale and advanced features at the cost of operational complexity; PostgreSQL full-text search provides language-aware normalization and ranking for longer text, while trigrams are language-independent and tolerant of misspellings but less suitable for large text. The transcript ends as he begins explaining how strings are split into overlapping three-character sequences.

Key takeaways

  • Django’s `contains`, `icontains`, and PostgreSQL’s `unaccent` lookup provide simple, fast matching but do not rank results.
  • External search systems offer scalability, spell checking, faceting, weighting, and advanced ranking, but require a separate index and database synchronization.
  • PostgreSQL full-text search normalizes words, removes stop words, supports language-specific dictionaries, and ranks matches across fields.
  • PostgreSQL full-text search works well for long articles in one language but can perform poorly on short titles and unusual names.
  • Trigram search is language-independent and can find misspellings, although it is less effective for large bodies of text.

Summarised automatically from the transcript.

Transcript

3,601 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

5:31

Speaker 1: Hello everyone, welcome to the Django Meetup. Marco, are you there?

5:37

Speaker 2: Uh yes I am. Hello. Oh still have my old uh my old background.

5:48

Speaker 1: Perfect. Well, welcome everyone to the August virtual edition of the London Django Meetup. This is our fourth virtual meetup now, isn't it, Marco?

5:58

Speaker 2: Uh I think the fifth maybe.

6:00

Speaker 1: The fifth?

6:02

Speaker 2: Yeah, it could be. I think we started in April, so from May Ginger. Yeah. I think it's the fifth

6:10

Speaker 1: Well, that's 2020 for you, isn't it? It's time going very quickly. All right, I'll just go through our welcome slides then. Welcome everyone. Thank you for joining us on this heat wave day. I'm not actually in London, so I'm not having the joy of experiencing it. I'm over in Lisbon. Okay, here's our code of conduct. Be awesome to everyone. This especially applies when we open up the video chat for everyone at the end of the meetup to carry on discussing Jenko. If you've not heard of our code of contact, you can go check it out on our website with the link there. And tonight's meetup is again kindly sponsored by these two organizations, JetBrains and Pollen.

6:58

Speaker 1: Poland is a marketplace for the best in shared experiences. And I should have asked if anyone from Poland is on the line. I'm having a quick look. I'm not seeing anyone I recognize. And if if you are from Poland and want to flag this, just ping on the chat. Okay, I think I think we're We'll carry on.

7:21

Speaker 2: Give them a minute to

7:26

Speaker 1: colon. co. I'm not sure if they're hiring right now, but they have been in the past. And they probably will be on the future. All right, I'll carry on. This is our schedule today. We've got a little bit of news. We've got a talk from Marco here. We've got a prize draw and then we'll open up a discussion. If you want to be promoted, you'll be free to turn on your webcams and audio and chat about whatever topics you feel are relevant. Perhaps asking more on the talk or other news that you've heard of. And so here's the big piece of news. Django 3. 1 has been released finally.

8:11

Speaker 1: The code name for this release is Poppuri. It's not really an official codename thing we do, but in every single release blog post there's an adjective used to describe the new features like a host of new features or in this one it's a pot parie new features. The two headline features would be the asynchronous views and a universal JSON field. So you no longer need a third-party package to do that. Other news is a couple of virtual Django conferences. The first one being DjangoCon Australia. That's Australia with an A in the middle, not Australia, as I've managed to typo there. And that will be online on the 4th of September. It's actually part of PyCon Australia. And I think that's free this year.

8:58

Speaker 1: I couldn't see any pricing information anywhere on their website But at the same time, it's all being put together last minute, so I don't know if they've shared anything about attending yet. And then we have DjangoCon Europe as well, which is online a couple weeks later, from the 18th to the 19th September. I know one or two people here, including myself, are giving talks here. I guess though it's mostly going to be recorded talks that will be played back online at the same time and ability to socialize. This is um Also free. This is this evening's talk. It's going to be on full text search in Django. And Marco's going to talk us through a couple of different approaches that he's used

9:46

Speaker 1: I think centering on Postgres at the end, but we'll find out more in a second. We also have a prize draw for two. discount codes for a one-year license of any JetBrains product. This is their sponsorship to us. It's very generous. During Marco's talk I will paste a link. To a Google form and I'll paste it again at the end. And then from that, we do a random draw of uh two people to get their licenses. And as I hinted at earlier, the end will be an open discussion for anyone who wants to stay on, discuss anything related to Chango. If you've got a burning topic or question, feel free to ask it to our Not so large room of experts

10:32

Speaker 1: this time. We'll have 23 people at the moment, it looks like. So it should be a nice size for a discussion.

10:40

Speaker 2: Um sorry Adam, there is a message from uh Benjamin from Django Day Copenhagen. He was saying that uh they will also stream on September 25th. In fact, actually like the talk that I'm giving tonight, it will be given again at January Copenhagen. Ben, do you want to announce it yourself since you are here?

11:01

Speaker 1: I'll promote Ben to a panelist. If you want to.

11:11

Speaker 2: Oh, we can do this later.

11:13

Speaker 3: Hello everybody. Yes, it's uh it's uh official news. Uh unfortunately we haven't been able to put it on the website yet. Uh it's Tonight we're learning a little bit about uh conducting meetings on Zoom. And uh we are confident now that we have both uh plan A and a plan B for having both uh physical and uh an online virtual Django Day Copenhagen 2020. So uh looking forward to your talk Marco.

11:43

Speaker 2: No, thank you.

11:44

Speaker 3: Both here and uh in a month.

11:48

Speaker 1: Awesome and awesome background.

11:53

Speaker 2: Oh nice

11:57

Speaker 1: I guess that's from DjangoCon Europe last year.

11:59

Speaker 3: It is indeed.

12:02

Speaker 1: All right. Thank you for that extra news. Is that going to be free and online, Ben?

12:08

Speaker 3: Yes, it is. And uh there would be a special package available for uh both people who uh In case we need to cancel, we'll simply convert all of the physical tickets to a package that's sent to the doorstep and uh it's also possible. very soon to to purchase a home viewing sort of experience for the for the day.

12:32

Speaker 1: Cool. Awesome. Thank you. Yep, so that would be the discussion at the end. One last point before we get on. If you want to give a talk, please email or message us. In fact, we have a form now that you can fill in, Google form. from which we'll contact you. And that is in a pinned tweet on our Twitter profile. So if you go to twitter. com slash Django London, it's pretty easy to find.

13:00

Speaker 2: I will also post a link later just for every word.

13:05

Speaker 1: Excellent. And here are our links, DjangoLandon. com. Twitter, GitHub, and Open Collective if you want to look inside our finances. And now we're going on with the talk.

13:22

Speaker 2: So yeah, hello everyone. Um let me figure out my screen. Um Uh Adam sorry but you have disabled this share screen for me

13:50

Speaker 1: Try now.

13:55

Speaker 2: Uh yes, appear it is Working. Uh can you all see my screen?

14:03

Speaker 1: I can see it and I will turn myself up there.

14:07

Speaker 2: Okay, thank you Adam. So full text search in Django. Um why do we um Why is SERS so important? Why are we talking about that tonight? Well, we we live in a world of data. Every day we are publishing new information, we categorize it, we share it, but we also need to access this. So we are different ways to categorize data. We have vertical system with categories. We have horizontal system like tax. But Categorizing formation is a slow task, is a long and slow task that is subject to systematic error and bias. There are some new machine learning systems that allow us to do this automatically and very

14:55

Speaker 2: fast, but setting them up is very time-consuming and expensive. Often search is still the best solution. You uh you just insert a keyword into a system and you end up with like seeing all the matches. is something that we have been using for years. User know this approach. So tonight I'm gonna talk about how to you implement this in Django. I'm gonna get through different technologies that are available in Django to do a full-text search and try to pinpoint what is the strong suit of each one and so when you should use each one based on your requirements. I divided the

15:41

Speaker 2: different technologies for search in four main families. Standard text or queries, external services. The PostgreSQL full text search and the PostSearch trigonom. The Sun Text query is the simplest search system that we had. Required zero setup, it comes with every database and uh is quick and simple to use On the other hand, it's a binary system. The only thing is that it's going to tell you if it's a string it in a specific record or it's not. will not tell you anything about how many times in that record, in what position, and it will not allow you to rank the result.

16:30

Speaker 2: Django comes with three lookup that belong to this family of the standard text and queries. The first two contain and icontains are available for every database. First one it will check for a string and it's case sensitive. It will make a distinction between uppercase and lowercase characters The second one, the I stand for insensitive and it will search for both the uppercase and the lowercase version and give you both results. There is a third one, an accent, that is Postgres specific. It will search for case insensitive, but it will also replace every accented character with the non-accented version.

17:15

Speaker 2: In this way we'll be able to match both. So if you look at this example, I have a Django using some specific accented character This kind of search will match also the non-accent version SSA standard query, standard text query is easiest and fastest way to do a text search. It has what 's a good case for this kind of queries? I think that filtering in the admin interface is an uh an excellent case. We are talking about a situation in which you have an internal user, not an external one. So someone that knows the system and someone that probably knows the

18:02

Speaker 2: data set already. So even if the search is a little bit is not that precise, they can still use it in a very efficient way. On the other hand, I will not use this for an external user. So if I have a blog and I want to search in the content of every blog post, I will not use this kind of search for the end user. is just not precise enough. Next on the list, external services. Now this is a big bucket in which I could a lot of different things. I mentioned three here, but there are many. Elasticsearch, Apache Solar, and Algolia. Also in these three

18:49

Speaker 2: they are kind of different because while Elasticsearch and Solar they are software that you can install and deploy by yourself, Algolia is It's a service that you can pay, they have their own API and you can use it, but you don't have to deploy it. This system are the are extremely performant. They can manage millions of record and they are particularly reliant to having like a lot of requests at the same time. And they come with a lot of features, you know, very specific to search. I'm talking about spell checking, faceting, weighted search, special ranking and more and more.

19:35

Speaker 2: It's almost too good to be true. On the other hand, they do increase the complexity of your application because now we have you have another piece. And this means that You have another component that is in production, probably if you have a signal environment, it's there as well. If you develop on your personal system, you also have to manage that locally And you have to manage two copies on your data because you're gonna have data on your database as usual for every functionality, and then you're gonna do search on a different system And so you have to take this two system and say. Now, external services is like the ICO

20:21

Speaker 2: side value solution. They are a little bit hard to implement, but They offer the highest performance and the most feature. How does the implementation usually look like? uh beside this being uh a lot of different system is it's usually more or less the same. Uh you will I have to define a schema for the search index that is kind of a model definition of this of this index for the search. You will have to index all the existing record in your database. So you are gonna have an asynchronous task or something that is sending to this third-party system. Everything you want to search for, and then you will have to find a way to synchronize between

21:10

Speaker 2: your database and the search index This is usually done through post save and post-delete signals in Django and If you don't know our signal words , I can talk a little bit about them at the end of the talk. But yeah, the idea is that every time that you change something, Jangwe make a post-save signals, and this will allow you to. index the new record in the search index and when you delete something Django emits a post -elist signals and you can use the signal to remove something from the index and in this way you can keep Your index up to date.

21:57

Speaker 2: What are this system good at? As I say, very big database, but also high traffic uh these are like um a system that are like bottle tested to manage a lot of requests at the same time Also, they don't necessarily have to connect to Django. So if you have a front-end application, you don't need to have uh an endpoint in the middle. You can have your front end talking directly with the search system and this is gonna save your request in the back end if you have an eye load. uh also some of the features some of the advanced features it could be very interesting for you uh for instance once i worked with um

22:43

Speaker 2: um a dollar based of books and one of the requirements was to having the book that um have been published later to come highest in the search than the book that we published earlier but with the same uh with the same title so if there is a new edition of a book that would come first in the search that's something that you can do with this kind of systems uh and provide a specific um function for ranking. One when you should know. to use this kind of system where you shouldn't avoid them. Well um as I said uh it requires some work to put them together. Um They come sometimes with Django-specific integration or

23:28

Speaker 2: Python integration that you can use, but there is always a little bit of work to integrate them and to maintain them. If you are tied on development resource, you should probably avoid them Next on is Postgres full text module. In this talk there are like a lot of very specific Postgres features. I guess that it is the database that is more supported by Django and I always suggest to go for Postgres when you are in a new Django project. So how does the full text search work? It's kind of a semantic search.

24:13

Speaker 2: I wrote ish in the slide. It's not very semantic, but it kind of used dictionaries in different languages too. kind of create um a semantic search system. It can use the advanced ranking. It has some features for ranking different results that are very similar to external system like uh elastic and solar. You can search in more than one field at the same time and it's highly configurable. On the other hand, is language dependent. And when I talk about language, I'm not talking about Python, I'm talking about English or French, Italian. It used dictionaries or word to work and so every time you index something you have to specify which language is

25:03

Speaker 2: it. But how does it work in Kanawa? So when you want to index something, first it will try to normalize the word. So take a word and remove the plural, for instance. Then we'll remove from the text that is ready for search every stop word like preposition and conjunction and in general very common word that will create noise in the index. And in the end, it will apply to every word and a frequency score and use them in ordering. Let's do an example. I use it The Postgres SQL query here in the example and also in the following, but Django does this for you

25:51

Speaker 2: underdue with your RAM. So if you try to index the Django tag line, the web where uh the web framework for perfectionists with deadlines, uh you specify the language English. First is going to normalize the word. As you can see, that line, that lines, has become that line and the last part of the word has been removed. So this can be matched with both singular and plural. Uh D, four, and width have been removed as very common words. And eventually we have a frequency score applied so deadline is not a very common word and as a seven um

26:36

Speaker 2: while web is a more common expression instead just a skew and framework a perfectionist over in the middle Um so wait a second. Oh let's keep on this page. Um Postgre applied this algorithm to both everything you want to search in, to index it, but also to the query that you're giving to them. So when you try to search for something, it will apply the same algorithm and then compare the two result and then give you um give you a result. It's interesting to notice that uh this This normalization algorithm is language dependent.

27:23

Speaker 2: And so if I take the exact same sentence and I ask Postgres SQL to index it for different languages, the result will be different. In Italian, for instance, the word for uh will not be filtered out because it's not a stop word in Italian. As deadline, it would be not normalized because uh singular and plural rules they work in a different way. This can be a limitation. If you have a data set or with records in few different languages, the search is not gonna be as good. as for dataset for which uh every single record isn't the same language

28:10

Speaker 2: Another very interesting feature of the Postgres full text search is the ranking system. It will give you highest ranking when you match multiple words in the same record. If the two words that you are matching, they appear together in the same part of the text. They will give you an IR rank that if one appears at the beginning of the text and the other one at the end. And also since the system can search in different fields at the same time you can apply a different weight, different importance to each individual field. So a match in the title will be ranked higher compare it to um a match in the description or

28:57

Speaker 2: in the content. Why you should use this? I think that like for uh content of the articles of a blog or uh a magazine, this is a very good example. This works very well. You have a lot of long pieces of text in the same language and you want to find sparse information in them. A better use case will be movie titles This is a very specific example. I want to explain this. Movie titles often they use some weird spelling or they use noun in them or non -common words. So normalization it

29:43

Speaker 2: will not work on this. Um sometimes the The article is a very important part of the movie title. So like Lord of the Rings, it will just be Lord Ring. And you 're removing half of the words. So yeah, normalization for really short pieces of text that have a night density of noun, it can be really bad. And it could actually work against me. Yeah, last but not the least, actually my favorite Trigram. Trigram is language independent, found resolved, even when you have misspelling. On the other hand, is not as good as

30:30

Speaker 2: full-text search for um for large text sizes. So our program works. Um, the algorithm will divide every piece of text in three grams. A trigram is a series of three characters. This is then update first Some space at the beginning and some space at the end of every string. Then the system goes and takes every at the first three characters Then ignore the first and take from the second to the fourth, then from the third to the fifth, and

31:17

Speaker 2: create all this series of three characters. In the end When I transformate a piece of text in a list of trigrams, remove duplication from this list and order them. I know this cancel are complicated, but I'm gonna make an example. So the string jangle, for instance.

Questions this talk answers

What are the pros and cons of using Elasticsearch, Solr, or Algolia with Django?

External search services provide high performance, scalability, and advanced features such as spell checking, faceting, weighted search, and custom ranking. They also add operational complexity: you must manage another component, duplicate search data, and keep the index synchronized with the database.

Discussed at 18:49

How do I keep an external search index synchronized with Django?

Define a search-index schema, index existing database records through an asynchronous task, and use Django's post-save and post-delete signals to add, update, or remove indexed records as the database changes.

Discussed at 20:21

How does PostgreSQL full-text search work in Django?

PostgreSQL normalizes indexed text and search queries using language-specific dictionaries, removes common stop words, and applies frequency scores for ranking. Django provides the corresponding functionality through its PostgreSQL search features, so the database compares the processed query with the processed indexed text.

Discussed at 24:13

What are the best use cases and limitations of PostgreSQL full-text search?

It works well for long articles, blog posts, and magazine content where users need to find sparse information in substantial text, especially when ranking and searching across multiple weighted fields matter. It is language-dependent and can perform poorly for short titles or text with unusual names, spelling, or important stop words, such as movie titles.

Discussed at 28:57

What is trigram search in PostgreSQL, and when should I use it?

Trigram search splits text into overlapping groups of three characters, making it language-independent and useful for finding misspellings. It is less suitable than full-text search for large bodies of text.

Discussed at 29:43

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Ashley Grealish and Marco Alabruzzo

More videos from Django London