Tips and tricks for optimizing Django response times with Carmela Beiro

This video features Carmela Beiro at DjangoCon US 2022 in San Diego, California, USA.

Tips and tricks for optimizing Django response times with Carmela Beiro
0:23:20
Published November 3, 2022
1,824 views

After deploying a project in production and generating new data, it's common that some response time issues arise. We don't always code having this in mind since probably at the beginning of a project there won't be enough data to cause this concern. Therefore, it's important not only to take this into account when developing in order to prevent performance issues in the future, but also to be able to debug and detect the cause after they happen. After working in several projects over the years, I’ve found some good practices, tips and tricks that have helped me to prevent high response times or decrease them, making the process more efficient, and improving the overall user experience and satisfaction.

This talk was presented at: https://2022.djangocon.us/talks/tips-and-tricks-for-optimizing-django/

Follow DjangCon US 👇
https://twitter.com/djangocon

Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/

Summary

Carmela Beiro presents a practical process for diagnosing and reducing slow Django responses: profile the application, understand QuerySet evaluation and caching, optimize related-object loading, inspect database plans, and design indexes carefully. She explains how Django Debug Toolbar and Django Silk expose expensive SQL and Python code, why `select_related` and `prefetch_related` prevent excessive queries, and how index order, partial indexes, and selecting fewer fields affect PostgreSQL performance. When query and database tuning are not enough, she recommends caching, precomputing reports, returning results asynchronously, or considering a database better suited to the workload, while noting the trade-offs around stale data and index maintenance.

Key takeaways

  • Use Django Debug Toolbar in development or staging to find slow admin pages, SQL, signals, cache behavior, and Python code; use Django Silk to profile requests and integrations.
  • Remember that QuerySets are lazy, and order operations to avoid repeated evaluation; QuerySet caching does not generally apply when accessing individual indexes or slices.
  • Prefer database-level operations when appropriate, but avoid replacing a later full evaluation with `count()` if that causes the same data to be queried again.
  • Use `select_related` for foreign-key relationships and `prefetch_related` for many-to-many or separate related queries to avoid N+1 query patterns.
  • Design indexes around actual query patterns: column order matters, partial indexes can reduce size, and every index adds write and maintenance cost.
  • For remaining bottlenecks, cache results, precompute expensive reports, process work asynchronously, or choose a database suited to time-series or analytical workloads.

Summarised automatically from the transcript.

Transcript

3,658 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:21

Speaker 1: Hi , my name is Carmela Beiro. My Twitter handle is Carme Beiro. I will be posting this the slides afterwards. It's going to be my first Twitter post. I'm a data engineering at Octavo Dev, that is a software development consultancy. We are based in Montevideo, Uruguay. Um it's Uruguay is a small country in South America next to Argentina and Brazil. We are about uh three million people. It's a really small country. And well, some of my hobbies or passions are reading, uh animals, uh doing workouts such as yoga and traveling. Okay, well, uh it would be ideal to at the beginning of a project know all the response time issues that we are going to have, but we all know that that is not possible.

1:11

Speaker 1: Uh first because we cannot foresee every single uh issue that we are going to have, and second because we don't know how the data is going to be in production So the idea of this talk is to present um well what steps uh have I followed in the different projects where I had these issues and also what considerations to take. So, in order to present the different examples, I'm going to use these two models. We have recipes That have a name, instructions, they have an author, a category, the difficulty, and also if it's active or not. And then each author has a name and an email. Okay, so once we have an issue in production, what uh what can we do?

1:57

Speaker 1: Well, um we could use a profiling tool. The idea of these tools is to provide a way to know which part of the code is being executed, how much time it's taking to execute, and also the SQL that is being executed. So I'm going to present two tools that we used uh over the years. Probably there are more, but these are the ones that I've known. The first one is Django debug toolbar. This is used in Django admin. It appears like a sidebar. So you can see for each page that is loaded in the admin, which are the SQL code that is executed, if there are any signals that being are being executed as well. And uh well uh if the cache is being used and you can also profile Python code.

2:43

Speaker 1: This was really useful because we had a client that's our first That the first request was to analyze why their Django admin was taking too long to load when listing a certain amount of instances of a model. They had around uh eight hundred eight hundred thousand instances and it was taking approximately fifteen seconds to load So the first thing that we did was install Django debug toolbar to determine what was the cost of that bottleneck. So we realized that um they had a filter that was uh it was a multi-select filter that it was loading all the related instances to each of those um instances of the related model and that was uh taking a lot of time

3:29

Speaker 1: so we replaced that with a search box with an autocomplete to only load the ones that uh were filtered by that criteria and that helped a lot So I really recommend this because it was really useful for us. Another tool that we find useful is Django Silk. This is mainly used for requests, so we use it in our integrations with Django Rest framework. They provide an initial screen where you can see which are the requests that are taking longer. Which is the SQL that is being executed. And you can also see well um particular requests, the SQL, and you can also profile Python scripts. model

4:14

Speaker 1: functions as well, that you can analyze the dependencies that the code has and well which part is taking longer. So one note about these tools is that they are not meant to be used in production because they add an overhead when when or loading the admin or performing a request. So the idea is to use them either in development or in staging. Okay, once that we detected uh what the issue is, another important thing is understand how Django works behind the scene. Not only for fixing a bug, but also when development is good to have these things under consideration. So the first thing that I wanted to mention is that query sets are lazy.

5:02

Speaker 1: This means that you can change any operations in the query set such as filtering, excluding, uh union, distinct, or thereby. And those queries are not being are not going to be executed until you need the data, so until you need to evaluate that query set. When does that happen? Well, for example, when you need to loop through a query set, when you need to calculate the length, like I'm showing in the example. When you need to check um with a Boolean condition if the query set has any values inside. Or for example when you do slicing. In this example, I'm doing some sequential operations that are not triggering any database calls, but then when the length uh is calculated, a database query is performed.

5:49

Speaker 1: So until I don't need it, the the query is not performed. This is useful to know because if we order uh the operations correctly, we may avoid unnecessary database calls. Another important thing is that Django caches query sets and attributes. So when a query set needs to be evaluated, what Django does is it first checks if the cache already has the query set stored. If not, it performs the query and then stores the results in the cache. So then on the next evaluation, instead of querying the database, it uses the results from the previous cache. In this case, I'm performing the same operations. I'm checking the length that is going to check

6:35

Speaker 1: the database because I don't have the data already stored in the cache. It's going to store it and then when I'm going to perform the loop, um the data is already cached, so I don't have to perform another database query. Well, there are some cases where after evaluating a query set, the data is not stored in the cache When does it happen? When um I'm checking only a portion of the query set. For example, when I need to access an index a specific element or when I'm slicing. So in this case I switch the order of the length and the and the loop and since I'm performing a slicing in the recipes

7:20

Speaker 1: um that is queried but it's not cached. So there is one database call and then in the next one since I don't have the data in the cache The query is going to be performed again. So in this case, if the order is the same to us, it may be more performant to perform the length before the four. Another important thing here that I did for the example is I moved the order by below the lens. So in this case, I performed the length, the data is stored in the cache, and since I performed another query set operation, Django needs to query again the database for the loop. So this is important as well.

8:06

Speaker 1: Um for me the order byte doesn't affect the length. So it's better to define it at the top and not afterwards, so I can avoid another unnecessary database call. Okay. Since Django provides different ways of performing the same operations , it's important to determine which is the best way. Almost always the best option is to perform operations at the lowest possible level. In this case, between Django and Python, it sorry, between the database and Python, it's the database So in this example, I'm trying to determine if an author has recipes. In the first example, I'm checking with a Boolean condition if an author has recipes. This is using Python.

8:54

Speaker 1: In the second example, I'm using Exists that is going to query the database. So in these two cases, it's best to do the second option. But it's not that simple because if I know that I will be querying um evaluating the query set, it's better to do that um as soon as possible. So then I can I can perform Python operations instead of querying the database again. So to show you this, I created this example. Um where I changed the length for the count. So instead of doing it in Python, I'm doing it in the database. That is performing a query It's not storing it in the cache because I'm not querying the whole query set.

9:39

Speaker 1: And then when I do the four, I need to query the database again. So in this case I know that I'm going to evaluate the query set because I need it for the loop. So in that case it's better to do perform the lengt operation that is going to bring it into cache and then perform the loop and in that case I will be having only uh one query to the database instead of d two in this case. Another important aspect is how Django fetches related objects Uh this was mentioned um in other talks so I will be uh a bit uh I will pas this this a bit quicker but uh in Herdinger database queries and the Django admin is your oyster they mentioned uh how this works. So in this example, I want to get all the authors related to recipes.

10:28

Speaker 1: Let's see how uh Django handled this query. Well, so as you can see there are a lot of queries. Um the first one is selecting all the recipes And then there is one query per author per recipe. So I had 4,000 um recipes in my database, so this is going to perform 4,01 uh queries. So this could be pretty inefficient when we have a lot of data. So what can we do to avoid this? Well Django provides two methods, prefetch related and select related, that lets you pre-fetch related objects. So select related works for foreign uh foreign uh foreign key relationships.

11:14

Speaker 1: And uh prefet related for many too many. What select related does, it's performs a join between the two related objects and it performs only one query. Prefetch related performs two queries, wants to get all the recipes and another to get all the authors related to the recipes So in this case, I'm showing what's the query executed when I perform a select related with the author, and you can see that there is only one query that performs an inner join between the author and the recipes. So instead of doing 4001 queries, I'm doing only one by using select related. Okay, um once we had detected what the issue is, and we had determined that the order

12:00

Speaker 1: of the evaluations and the operations in the query sets are correct and we are doing everything in the most efficient way, we can check how the database is executing our queries. In particular if uh any indexes are being used or not and if we need to define new ones So uh in order to define an index, um the most important thing is to determine which are the queries that are going to be executed. So it's essential that the developers that work with the database are the ones that define this. Um after we define an index uh the query optimizer that is the one that decides how the query is going to be performed not necessarily is going to pick the index that we defined. And this could happen because of two different things.

12:46

Speaker 1: The first one is that the query optimizer uses um statistics to determine if it's um if it's more performant to select the index. But it could happen that uh our index is the best solution, but we haven't decided we haven't defined it in the correct way. So the optimizer is not picking it. And another important thing is that adding an index adds m maintenance costs. So every time we need to create, update, or delete An instance that is related to an index, we need to update that index. So those operations are going to take a bit more. Um so that's where the cost appears. So uh the best thing to do is to define the less uh amount of

13:31

Speaker 1: uh indexes to cover the greatest amount of queries So now I'm going to show you some things that it's good to have into consideration when defining indexes The first one is that index orders matters. So let's consider that I have a web app where my users need to select a category and then they have a search box for entering the name of the recipe. So when they enter the name I need to search uh all recipes that starts with that string. So in order to do that, I decide to create this index that has uh the name and the category. Since um the start with operation is performed with a like operator in Postgres, that is the database that I'm using, I need to define

14:20

Speaker 1: the op class for the name and say that the name is barchar pattern ops and that's the only way that Postgres is going to be able to perform like operations with the index. So let's see how the query performs. What I'm doing here is I'm filtering the recipes that start with soup and has the category breakfast and I'm explaining um I'm I 'm selecting explain that is an operation that Django provides to see what the query optimizer is doing. Here I can see that the index scan is being used, so that's great. Let's see what happens if now I decide that I only want to query per category. Well, um we can see that a sequential scan is being used, so the index it's not being used.

15:09

Speaker 1: What happens? that the index order matters. So if I define my index as name and category, I can only use the index for query name or name and category, but not category alone. It works like a phone a phone book. So you have order your phones by first name by last name and first name. You can query it by last name and you cannot query it but first name alone. So it's the the same um the same thing Okay, another thing that could be useful is defining partial index. So it's a way of adding a where condition in your index For a condition that is applied uh almost every uh every time when you are performing a query.

15:56

Speaker 1: So in this case, most of the time I want to know which are the queries, which are the recipes that are active. So What I did is uh Django provides an option for def defining a condition that is like a wear condition where I select all the recipes that are active. This could improve the performance of the queries because Uh my index is uh smaller because I I'm considering less rows, so this could be more performant. Another thing that is really important is trying to select only the attributes that I'm going to use. This is important because this is going to consume uh to get less data from the disk and also when performing operations I'm going to have to consider less data so

16:42

Speaker 1: probably it's going to be more performant. And the other thing is that um the probabilities for being able to use an index-only scan are bigger if I only use uh uh less attributes because If my index has the same attributes as my query, then the index can be used alone to get all the data and you wouldn't have to check the tables in the disk. So this could improve um everything the performance uh a lot. So Django provides two ways of selecting only a small portion of amounts in a table. The first one is values where you indicate the attributes that you want to lose to use and it returns a query set with dictionaries inside. And then the other two options are using only

17:28

Speaker 1: and defer only selects the attributes that you want to query and defer the ones that you don't want to query. And it returns a query set with less amount of attributes. So then you have these two options, you can di decide which is the one that uh fits best your your requirements. Okay. What happens if um I checked everything and I still need to improve the performance? Well What we have done in the past is first of all using caching. Django provides uh a the corrector for functions um for caching the results. So when the operation needs to be performed, Django first checks if the results are already stored in the cache. These stored results are going to be valid for a certain amount of time.

18:16

Speaker 1: And once they expire, they are going to be removed from the cache. So if they are not stored, the the function is uh performed, the results are stored, and the next time I perform the function, maybe I can use the results that are already cached. This is going to improve the response time, but it's going to have a trade-off that is that the results are not going to be valid for a certain amount of time. The amount of time that I define for the results to be valid. So that's something to consider. Another thing that we found useful is to define is to store pre-calculated results. So for example, if we have a report that um has to do some complex calculations and has to do some aggregations, then what we have done is

19:01

Speaker 1: with a salary job that uh runs on a schedule We calculate the results and we store it either on the database or maybe in in the memory using memcache or Redis, and we already have that available when the user needs to query it. Again, there is a trade-off that is that the results are not going to be valid, uh not valid, but the results are not going to be up to date between the period of time that the job runs. Another option that did dep this depends on the requirements of my project is um I could return uh results in an asynchronous manner. So for example, I could say to the user, okay, I'm calculating this, I would send you an email once the results are ready.

19:47

Speaker 1: And the other option is well I would send you a notification. Again, this is not always possible because of the requirements in our projects. Well, and then another thing, another important thing, is that maybe the type of database that we are using is not is not the best one for um our requirements. So for example If I'm only all if if if I'm always squaring um data related to time and I need to perform calculation based on time, maybe a time series database is more useful Than a relational one. Or for example, if I need to query a lot of historic data, perform analysis and aggregations, a columnar database is more useful. So

20:33

Speaker 1: One uh one needs to check the requirements and see which is the database that um performs the best in these scenarios. So here there are some useful links. The Django documentation that is very good that lists all the steps that I mentioned and explains how the query sets work behind the scenes the two um tools that I presented Django debug toolbar and Django Silk and last there is a book that explains in uh indexes And all the things that I mentioned and more, uh that is really good, that it's uh nice for checking. Thank you.

21:16

Speaker 2: Yeah, I had a question about the uh setup uh for when you would uh not want to use the async. Um the only thing I can think of is if um If you're not using ASCII, right? Um and if you're not using ASCII, then I'm thinking you can't use the async approach to fetching the data.

21:39

Speaker 1: So what happens if you cannot use an asynchronous approach?

21:46

Speaker 2: Yeah, I mean the the question is yeah i i i if the um the requirement is to want to async fetch your data, right? Then I guess what are the requirements so that you can um be guaranteed that you can use the async uh when retrieving the data?

22:06

Speaker 1: Uh n the requirements in Django or in general?

22:10

Speaker 2: Um I'm decided on using Django. Yes. So uh based off of that, yeah. Yeah.

22:16

Speaker 1: Okay, so you should return a response if you are uh return if you are for example using a rest uh framework, you choose return a response. And once uh you can use celery for example to trigger a job to run the results and then when that is finished for example send an email with the with the results. You could do that

22:37

Speaker 3: Does um prefetch related and select uh related work if I understand if it's a you know relationship, but what if they're chaining things, you know, re you know the

22:51

Speaker 1: More than one relationship you can pre-fetch uh the other related uh object. Oh really?

22:56

Speaker 3: So we can do so it doesn't it doesn't have like a uh the depth level of how low or

23:02

Speaker 1: I'm not sure. I I I'm not sure. Maybe yes, but I think you can change uh at least two for example.

23:10

Speaker 3: Okay, thank you

Questions this talk answers

How can I find what is slowing down a Django page or API request?

Use a profiling tool such as Django Debug Toolbar for admin pages or Django Silk for requests. They show executed SQL, timings, and Python code that may be responsible for the bottleneck.

Discussed at 1:57

Can I use Django Debug Toolbar or Django Silk in production?

They are intended for development or staging because they add overhead while loading the admin or handling requests, so they should not normally be enabled in production.

Discussed at 4:14

When does Django actually execute a QuerySet?

QuerySets are lazy: filtering, excluding, ordering, unions, and similar operations do not hit the database until the QuerySet is evaluated. Evaluation happens when you iterate, calculate its length, test it as a Boolean, or take a slice, among other operations.

Discussed at 5:02

How does Django cache QuerySet results?

After a complete QuerySet is evaluated, Django stores its results in the QuerySet cache and can reuse them on a later evaluation. Partial evaluations, such as indexing or slicing, generally are not cached in the same way and may cause another database query.

Discussed at 5:49

Should I use exists, count, or len when checking a Django QuerySet?

The best choice depends on whether the QuerySet will later be evaluated. If you will iterate over it, using `len()` can populate the cache and avoid a second query, whereas `count()` queries the database separately and does not cache the full results; for a simple existence check, a database-level `exists()` is preferable.

Discussed at 8:54

What should I consider before adding a database index to a Django model?

Start with the queries the application actually runs, and remember that the database optimizer may choose not to use an index. Indexes also add write and maintenance costs, so use as few as possible while covering the important queries.

Discussed at 12:00

Why does the order of columns in a database index matter?

A multi-column index can be used for the leading column or for the leading columns together, but not generally for a later column by itself. An index on `name, category` can support searches by name or by both fields, but not category alone.

Discussed at 15:09

What is a partial index and when should I use one in Django?

A partial index includes only rows matching a condition that is commonly used in queries, such as active recipes. Because it contains fewer rows, it can be smaller and faster than a full index.

Discussed at 15:56

How can I make Django queries faster by selecting fewer fields?

Select only the attributes the application needs, using `values()` for dictionaries or `only()` and `defer()` on model QuerySets. This reduces data read from disk and can make an index-only scan possible.

Discussed at 16:42

What can I do if Django queries are still too slow after optimizing them?

Cache function results, precompute expensive reports on a schedule, or return long-running results asynchronously through a job system and notify the user when they are ready. Depending on the workload, choosing a more suitable database—such as a time-series or columnar database—may also help.

Discussed at 17:28

How can I run a long Django calculation asynchronously and notify the user?

Return an immediate response, trigger a background job such as a Celery task to calculate the result, and then send the result by email or another notification when the job finishes.

Discussed at 22:16

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon US