Django from queryset to serialization with Iuri de Silvio

This video features Iuri de Silvio at DjangoCon US 2022 in San Diego, California, USA.

Django from queryset to serialization with Iuri de Silvio
0:23:54
Published November 3, 2022
790 views

A common Django project bad practice is serializing ORM objects without properly loading relationships, causing N+1 query issues. Django don´t have an obvious way to avoid that, but some techniques and libs can help to produce better code without too many unnecessary queries.

This talk was presented at: https://2022.djangocon.us/talks/django-from-queryset-to-serialization/

LINKS:
Follow Iuri de Silvio 👇
On Twitter: https://twitter.com/iurisilvio

Follow DjangCon US 👇
https://twitter.com/djangocon

Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/

Summary

Django applications often turn model instances into JSON through ad hoc serializer methods, but adding related fields can quietly introduce N+1 queries, expose unnecessary data, and make query optimization the responsibility of each view. Iuri de Silvio argues that data loading and serialization should be defined together: serializers should declare the fields and relationships they need, while custom query-set and manager logic prepares that data efficiently. He presents Django Q Serializer, a library that follows this pattern, supports bulk loading of external data, and makes query counts easy to test; he also explains why DRF and automatic prefetching do not fit his preferred approach.

Key takeaways

  • Serializing related Django objects without planning the query can create N+1 queries that become costly as traffic and data volume grow.
  • Using select_related or prefetch_related fixes individual cases, but keeping query logic in views and serialization logic elsewhere is difficult to maintain.
  • Each serializer should declare the data it needs so a custom manager or queryset can load it efficiently before serialization.
  • Django Q Serializer applies this pattern, including hooks for fetching external data such as bus locations in bulk.
  • Testing serializers with query-count assertions helps prevent later changes from accidentally adding database queries.
  • The speaker prefers explicit loading requirements over automatic prefetching because explicit code is easier to review and understand.

Summarised automatically from the transcript.

Transcript

2,726 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:20

Speaker 1: Hi. First, talking about me, I'm Yuri, I'm a principal engineer at Boozer. Boozer is a Brazilian company, so we have some buses, pink buses. And I'm saying that because most of my examples are about buses. So it's easy for. And I'm from Brazil Oh how we started with uh yeah I will talk about my experience at at Boozer and about another companies and other another projects and most of them start the same way when you have a model

1:06

Speaker 1: You have to convert convert it to uh a dictionary to to serialize it to pass through HTTP or something else. So at the end of the day it's just a JSON. So we have to Convert it and there's no not an a simple way to do that. So most projects start with a simple method to just copy some attributes to a dictionary so your view can just call it And it works. It's a simple view. It's a JSON response.

1:52

Speaker 1: I list all my buses here. They are really simple. They just have their plates So it works but when you we do that uh we start to add more attributes to our ser serializer this function so We we start with a a we start creating a problem here because now I need to add an owner to my bus. So I I have to serialize that and when I do that with a foreign key, when I call the what Self

2:37

Speaker 1: honor uh it's a cu a a curry and And it's not explicit to a new developer, uh someone that don't know Django, and it still works. So now you have that. We have uh The same view with the same method, but now we have n plus one queries, and that's the problem that I had most of the time with Django and with most RIMs uh it's not only Django fault so we have that and There's not an

3:22

Speaker 1: easy way to fix. So I started trying things. That example is when I I joined the company, Boozer. It's it was a it still is a fast-growing project. At the time we are sponsoring a soccer team We are doing TV ads and it was November so we have had uh Black Friday there with crazy uh promotions and our pro the project was already big so it was a we each model had a serialized method like that Not like that with a lot of relationships. So

4:08

Speaker 1: our servers and our database were always overloaded, and my main job there was to start improving their their code base and that's the thing I'm doing since the beginning three years ago and it's an endless job and I really like that the subject. So how to fix it? Uh here we had here we add a did uh owner And now I have to make sure that uh I have only one query or something better than any plus one curious

4:53

Speaker 1: queries because when I I have Just some objects you don't even know it's a problem, but when you scale it's a really bad problem. And or with a huge uh volume of requests you will understand how i it works. So It's easy to fix. You can just change our key reset to add a selector related and that way uh jungle will do uh join at the x sql sql level and now I can have only one carry again But it's not good

5:38

Speaker 1: because uh my query set is in my view and my serialization is defined somewhere else in the model so it's not easy to handle it and when it's just a an example but record base heads hundreds of views probably at that time and now a lot more so it's not easy to scale that code device so Django has Houston Care sets where I can encapsulate this issue and create my Trans transform that in a two-step problem. First I have to say what I want to serialize

6:25

Speaker 1: What I want to load, that's the owner, and after that I have to say that I want to serialize. So now I It's in my model. So here is the custom care set. And here is a uh uh I'm overriding the the default model manager. So now I uh in my view I can just call it to serialize and after that I can serialize and I have only one query again It works, it's better than the other solutions, but we with the solution We have only one serializer

7:12

Speaker 1: for all my model and each view needs something different, needs More attributes or less attributes. In some cases, if I pass everything, I can expose sensitive data. So It's it's still not good, it's better than before. Uh and because I have only one way to serialize, I will put there everything I need. So if I have 20 relationships, I will put everything there. But I need I probably never need all the 20. I need one in one view, another in other view, and if I

7:59

Speaker 1: I I don't have the option to reduce the problem I just put everything there So, what we're we are seeing there, each view has different requirements And my serialization m must know what it needs to load. And that's all context-dependent. So I have three steps. The first one I'm loading the data. I'm doing some business logic things here, and after that I serialize my data to send back to the clients. And we when we are cu creating our query sets, we are thinking about the first step.

8:47

Speaker 1: We load the data we need to Do our business logic, but we forget the next part that serial the serialization needs to know The serialization needs data that were loaded before, or it will go to any plus one problems or to uh uh bad ways to load this this data and it's not good. The and that that's the reason I I'm trying to

9:32

Speaker 1: do that. I'm trying to link loaded data with serialized data And I because of that I I created the uh a library. It's Django Q serializer. You can it's open source, you can install and check the GitHub that and how it works. It's I I just first override my my manager with a custom manager that's my library will handle And after that I I have to define serializers that know what they need to load.

10:18

Speaker 1: So in this case I know I'm defining the serialize object function, that's the base function for the serializer. That needs the plates and that needs the owner. So I can say for my serializer that you have to load select related and following the the past examples you can Just extend how it worked and instead of calling to serialize with nothing, now you can call it with your custom serializer And that way the the curie will will join with the company to

11:06

Speaker 1: load the owner and when you use the realizer you have only one uh query that way we if we almost fix the problem if I need uh each view needs a new a different serializer and I I can just create many serializers. Of course you need to create it thinking about not If you create many serializers, it's difficult to maintain. So it it's a bit catchy, but it's better than before. So here I have some examples

11:51

Speaker 1: and more um complicated than the base one. Here I I'm changing my key reset so I can use everything Django already has and just return I I just overhide the prepare carry set and return with everything I need and when I I'm going to serialize I just use the defined state. I'm annotating the the objects here. So I I I need that on in this case so it's all encapsulated and I can

12:36

Speaker 1: I I don't need to watch to see another file and another view and I know everything about the serialization based on on that Another example when I need external data, it's a common use case we have. We In in this case I I want to get my location of the buses. So I have a bus location service that doesn't matter what it does, but what matters is When I I I can override the process objects

13:21

Speaker 1: that receives all the objects in the this case are buses. So I get the geolocation for all my buses in an optimized way, and after that I I just attach it to my my object so now when I will serialize that I have access to my bus geolocation and I can do that and I know how I I access the data before and how to attach it when I have to serialize. After I built that I learned that it's really good to

14:08

Speaker 1: Test how much many carries I'm doing and doing that I I guarantee that someone no never will add another carry by accident and that's common. So Now I I can my it's uh just a uh simple jungle test. I'm I'm creating my bus here, my object. And s with Baker, so I just tell him to fill all all optional fields and uh relationships so probably I I'm covering most cases here just doing that and I can Test my serializer

14:54

Speaker 1: here just asserting that that we 'll do only one query, the only query that it does is loading the buzzes inside here the the the decorator the context manager and I know it it's uh I I know it it has only one query and it's very good because I can do that and I can do that easily for almost m all my serializers and my models in an easy way maybe it it can it can stay in the inside the library because it's really useful test to do

15:39

Speaker 1: This lib is I I like it to create this lib because it's easier to learn when you start using it. It's a bit complicated than uh change your queue set but when we you you have to do that you are forced in the right path you will always see what you have to load and what you can consume later and it's easier to code review too because everything is together and it's easier to test because you can count and test only the serializer not thinking about other parts of your system. So

16:24

Speaker 1: Keep counting is painful without this isolation. But uh it's not about this leap, my talk, it's more about the pattern that we found in a lot of projects. So that happens because jungle don't have an easy way and an obvious way like Zen of Python 6 and uh at Django Kong we had we already had a lot of talks talking about workarounds uh of R R or R M because it's a common issue. We we n need talks about

17:10

Speaker 1: select related, perfet related, how to handle that with jung DRF. uh that have serializers and other things and we have other people working things near what I did. Most of these projects are new, so when I started I already knew DRF but I don't use that and it has some issues. I already knew Django Auto Prefetch, but it doesn't solve the problem, it just hides it. You can't when you install AutoPrefetch. It will when you you execute a query set, it will

17:56

Speaker 1: automatically load all related objects with the same pattern. you don't know how it it's doing the their thing and it the problem here is you you you are not making the developer think about the problem so You're just hiding it in a better way than before, but you're still hiding it. And here I linked some projects. Django Readers, Django Strawberry Plus that had uh Tutorial yesterday, junk virtual models that Flavio will talk tomorrow, and probably he will talk more about all of them.

18:41

Speaker 1: So It's uh I issue that I want to fix or improve how we do that with Django because it's really painful to me and that's because I Talking about it. Thanks. That's it.

19:08

Speaker 2: Thank you very much, Yuri. We have time for a few questions. So if you've got a question, raise your hand. Um yeah.

19:15

Speaker 3: Thank you. Great talk. So I have a question. So I'm using because I can see in my scenario where I can fit this in my project is not having those query sets in the views and then I can move it easily to the serializer level I'm just wondering, does it affect performance? Have you compared the performance between the traditional way and then using this package? Is it gonna be decreased? Is it gonna be the same or is it gonna be more?

19:44

Speaker 1: I don't think it in it impacts performance. The the the serializer is really simple and it's it just calls some functions. You can see the implementation at the the repository but the implementation is really simple it doesn't don't have too much met too many magic too much magic it's just uh I I just organize the code Um

20:16

Speaker 4: I know this is uh an impossible question to answer, uh but What when would you say like uh because I feel like a lot of projects it depends on the scale of the traffic or the performance and like you know how do you know when to start thinking about this as a problem that you you would or is it part of your pitch is just use the library from the get-go and you'll you'll never even have known. Uh you know, like for when you're trying to quickly start with a project or on board, you know, how do you know when you've like crossed that level where it's important to Optimize.

20:52

Speaker 1: Yeah, I understand and I understand the question. Uh When you the project is small, you don't need to think about that. It will work for a long time and it just works. But When you start thinking about that soon you have the it's easier to scale because uh I I'm saying about Q serializer but still today three years later I started at the company we have a lot of serialized methods and it's impossible to uh uh remove them because it's a lot of legacy code. So if you start soon it will be easier to maintain the project.

21:40

Speaker 1: But I I always started thinking about that when it's already crashing. All my projects were this way. So I I I don't recommend that. We

21:57

Speaker 2: got another question. If there isn't, I'll slip one quick one in. Uh you uh you mentioned at the end the uh uh Django Rest framework and you said it had problems. Django Rest framework always would have been my sort of go-to for serialization. I'm intrigued what your what problems you were hit you were you were thinking against.

22:11

Speaker 1: Problem is you need to use DRF to use then. So it doesn't work does doesn't work for me. And I I don't like it's related to a view set, so it's still uh far from the serializer is far from the Q set so it's difficult to see when you change somewhere you have to change your QR set and It it it has this way the same problems at the end of the day. Sure.

22:44

Speaker 2: Any other questions? We have to do that.

22:53

Speaker 5: Thank you Yuri for the talk. Uh have you considered uh combining uh your serialization object uh method with some kind of tool like Bydantic? to format the data. For example, if you return a date time there, I guess you should format first to string, right? Yeah. So have you used those tools like Pydenic? I think it would be a great fit there

23:18

Speaker 1: Not yet, but it's a good idea. So maybe. Uh he is Flavio that we'll talk tomorrow.

23:35

Speaker 2: Okay, if there are no other questions, he thank you very much, Yuri, for your contribution to the conference and a small token of appreciation from the organisers.

Questions this talk answers

Why does serializing Django models lead to N+1 queries?

When a serializer accesses a related object such as a foreign key, Django may issue an additional query for each object. This is easy to miss in small datasets but becomes a serious performance problem at scale.

Discussed at 2:37

How do you avoid N+1 queries when serializing Django foreign keys?

Use `select_related()` for foreign-key relationships so Django performs a SQL join and loads the related data in one query. The talk also shows encapsulating that loading logic in a custom queryset or manager.

Discussed at 4:53

How can Django querysets and serializers share the data-loading requirements?

Define serializers that declare the fields and relationships they need, then have the serializer-specific queryset load those values before serialization. This keeps query optimization next to the serialization code and lets different views use different serializers without loading unnecessary or sensitive data.

Discussed at 9:32

How can you load external data efficiently before serializing Django objects?

Override the queryset processing step to fetch external information, such as bus locations, for all objects in an optimized operation. Attach the results to the objects so the serializer can use them without issuing extra per-object requests.

Discussed at 12:36

How do you test that a Django serializer makes only one database query?

Wrap serialization in Django’s query-counting test helper and assert that the expected number of queries is made. The example uses model factories to create realistic objects and verifies that serialization performs exactly one query.

Discussed at 14:08

Does django-q-serializer make serialization slower?

The speaker says it should not materially affect performance because the serializer is simple and mainly organizes calls without much magic or overhead.

Discussed at 19:44

When should you start optimizing Django queryset and serialization queries?

Small projects can work without this pattern for a long time, but adopting it early makes scaling and maintenance easier. Waiting until the application is already crashing leaves difficult legacy serialization code to replace.

Discussed at 20:52

Why not use Django REST framework for this serialization approach?

The speaker does not use Django REST framework because it requires adopting DRF and ties serialization to viewsets, leaving the serializer and queryset logic separated. That makes it harder to see and update both sides together.

Discussed at 22:11

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Iuri de Silvio

More videos from DjangoCon US