Operations: The Missing Django Piece with Micah Lyle
Published December 6, 2024
This video features Micah Lyle at DjangoCon US 2025 in Chicago, Illinois, USA.
This talk was presented at: https://2025.djangocon.us/talks/free-threaded-django/
LINKS:
Follow Micah Lyle 👇
Website: https://www.elyon.tech
Follow DjangoCon US 👇
https://fosstodon.org/@djangocon
https://x.com/djangocon
Follow DEFNA 👇
https://www.defna.org/
Video production by the presenter and DjangoCon US 2025 volunteers.
Free-threaded Python removes the Global Interpreter Lock (GIL), allowing multiple threads in one process to execute Python code on different CPU cores at the same time. Micah Lyle demonstrates this with Django’s development server, Gunicorn, and Granian: free threading can improve CPU-bound performance and reduce memory use compared with multiple processes, although results vary and some configurations use more memory. Django’s core is mostly thread-safe, but third-party packages and application code must be checked, since races can expose bugs such as work being performed twice. His recommendation is to remain cautious with synchronous Django, but to seriously consider free threading for CPU-bound ASGI applications once dependencies are compatible, measuring real workloads before switching.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Thanks so much for the intro. Hey everybody, I'm Micah. My wife and I run a small consulting firm called Elion Tech. And today I'm here to talk to you about free-threaded Django. on Wednesday at 3 p. m. at the end of the conference after a lot of our brains are fried. Here we go. So I'm going to give you an outline and give a roadmap of where we're going before we step into each part so that you can track big picture with all the things that we're going to cover. So, first, we're going to do a demo of running with the gill versus free-threaded and show some initial performance differences there to wet your appetite. Then we're going to talk about, well, what is the Gill?
Speaker 1: What's a thread? And we'll do a little crash course on that. And then we'll dive into is your Django code thread safe? And then if time, I think we'll have time, we'll fit in a little bit on async plus threading. So, how does threading impact your code if you have an async first Django app? And then I'll conclude with my opinion of the space and some recommendations and things to consider. So without further ado, let's get into this demo. Who here has used RunServer before? Raise your hand. Me too. I thought one of the funniest ways to start this would be to actually do run server for the demo. And the reason run server is really interesting with this demo, and
Speaker 1: I'm gonna go run this command. Is run server actually creates a thread for every single request? And you are not supposed to use it in production. I'm not gonna tell you if I have or not. But I've got this demo here, and this is gonna be a little hard to see, so maybe we'll zoom in a bit. And I'm going to hash a password using a uh educational purposes only Blake 3 password hasher. And I'm going to make a hundred requests all at the same time to my local laptop here, and you can watch the threads As they go, you can watch how long it takes for
Speaker 1: run server and my Python code to hash these passwords. And I've got another number of interesting things here. You can watch all the requests. This is as they're hitting the critical Django path, this waterfall, and I know it's a little tough to see. And basically what we're doing is hey , run server creates a thread for every request. Let's make it do some really CPU-intensive work, like using a pure Python hasher, which asterisk The uh most of the Python hashers actually run C code that release the GIL. So if you're trying to give a demo, use a pure Python hasher. Uh if you want to simulate like really heavy Python CPU work. And so I see a bunch of threads, I see some performance, and there's a certain number running at the same time.
Speaker 1: And that is my uh let's come back to the slides. Okay. So that's what we're doing. That demo's, I've made it open source, and so there's a number of ways you can run it. But some of the numbers I'm going to show are going to base off of that. And so these are things I just showed you, so I'm not gonna go over them in detail. But one of the things that was fascinating to me, and this graph may be a little hard to see, but was when I'm running Python with the Gill. I'm not seeing that much processor usage. Like maybe one spikes for a little bit, but in general, they're not moving a lot besides the one. Thanks to Richard and Nano Django, and thanks to UV, there was a pretty easy way to take this project and run it
Speaker 1: in the version of Python coming out next month. Uh in free-threaded mode. And I'm explicitly disabling the GIL because free-threaded Python is safe. And if there are third-party packages that haven't like opted into this new mode, it will turn it off. And so I'm saying no. Or it'll turn it back on. And I'm saying no, turn the gill off. And I'm not gonna go too much into the graph. There's some differences. But my memory usage monitors spiking on most of the cores now. So you can see they're all running at the same time. And to cut to the chase, Run server got forty percent faster on my laptop. Let's give it a round of applause.
Speaker 1: So if you run free threaded on your laptop, it'll probably be faster uh with run server. So at least as a developer You can have a faster run server. We'll talk about some of the uh side effects that come with that a little bit later. Um And the next thing I want to go over really quick is using that same demo, I was like, well, run server on my laptop is not going to be very helpful for folks. So Let's run G Unicorn on a bare metal server that I'm renting for the month that is doing nothing else but this. So we can maybe get some more reproducible performance. And The server did things a lot faster, and this is a 16-worker setup, and so G
Speaker 1: Unicorn, when you run it synchronously. Does 16 processes and one thread per process, and we'll talk about what that means in a little bit. And it had very predictable uh Performance here. You can see literally 16 of these cores hashing the password at the same time, then the next 16 are going, then the next 16 are going. It's very predictable. Very stable. Um and then I also ran it with this mode called preloading, which is where Unicorn loads your entire app into memory. Uh as best it can. It loads your whole app before it it forks out the child processes. Um and the performance is about the same, but I'm gonna show you memory usage in a little bit. The next thing I did, which is a really fun one to demo, is GUnicicorn has another worker type called G Thread.
Speaker 1: Some of you might use that in production, especially if you're using synchronous Django with a little more I. O. But that lets GUICORN say, hey, per worker process, give me a certain number of threads. But I said give me one worker, just one, and 16 threads, and let's see what happens Performance is very similar actually, but here you see the one process and the 16 threads. But it's running in free-threaded mode, so they're not contending for the gill, and the server waterfall looks very similar. As a bonus, there is a new kit on the block called Granian, and of course it's written in Rust, because what new stuff these days isn't. And I'm gonna put a disclaimer here.
Speaker 1: WSGI or synchronous, in this case synchronous Django, was added to Granian later. And free-threaded mode is experimental, and Granian has way more settings to tune. Than I'm used to. But I gave it a go, and Granian's chart's a lot more interesting. It uh splits the work up into threads. It seems like not as evenly, and there's a variety of of request and response times, and the waterfall looks a little funny. Um the funniest thing though, and you won't be able to see this, but I'm gonna tell you what it says. With Granian, we managed to accidentally make two ninjas. We created multiple ninjas and we crashed the program. So I was able to sometimes get this error and sometimes I didn't.
Speaker 1: And so we have some sort of thread race bug, our first one of the of the talk. And to compare everything together, it's gonna be a little tough to see at the bottom, but the Far right is grain with the threads. Maybe I had it configured improperly, but it was a little slower. The I'm not gonna look at total time because that counts request time, which is not There's some late latency variation there. But if we look at the performance comparison, GUICorn with the worker processes instead of free-threaded mode is a little bit faster. We're looking at maybe three close to 350 versus 380 milliseconds in this particular example. But there's not a huge difference there.
Speaker 1: Where the difference gets interesting, and sorry it's a little hard to see, bottom left is G Unicorn. Uh, the one after that is G Unicorn with a preloaded, and then the right two are threads. G-Unicorn running in G-threaded mode with just 16 threads has significantly less memory usage than if you're using a bunch of processes. And you would expect that because the operating system's doing a much heavier process. And so one of the potential benefits we're gonna see in this talk is with free threading, if the performance is about the same, And your code doesn't crash and there's not thread safety issues, you could potentially be a lot more memory efficient. Um, which I don't know about you all, but I've definitely rented my fair share of
Speaker 1: uh small instances on the cloud, and it definitely helps to be more memory efficient. Um all right. So now we're going to talk about what is free threading, and it's not that. And what's the gill? It's not that. And to answer these questions, well, we've got to talk about what's a thread. To know what a thread is, though, you should probably know what's a process. And with a process, we have to talk about operating systems. So we're gonna get a crash course in operating systems. Ten minutes, that's the goal. Don't time me. So we're gonna talk about from the inside out, we need to talk about the Linux kernel.
Speaker 1: And If you're already an expert in operating systems, you want to play with those demos. Ftdj. io hyperlinks you to the repository. But I think it's helpful for most of us to have a refresher on operating systems, especially given we're gonna be talking about threading and async and sync and all sorts of things. So Deep dive day. Here we go. I'm going to assume most of your production systems are running Linux and that the Linux kernel is going to be executing your code. If you're running a Raspberry Pi Windows server in your closet, this may not always apply to you. I have not met anybody that that does that yet. Alright, so we're going to start at the kernel layer.
Speaker 1: So your computer, when you turn it on, there is a program, or you could think of it as a piece of code running. that underlies all your other programs, such that when you type Python program and you hit enter, Python is being executed, but Python is just a program. There is something that is running your programs And also running most everything else on your computer, whether it's your Spotify, Apple Music, your 42 Chrome tabs, VS Code, whatever. And that kernel is also responsible for letting your program interact with the outside world. So if your program wants to print things out, if it wants to make network requests.
Speaker 1: If it wants to read and write files, which is typically how you get useful programs that interact with the outside world, your kernel mediates all of that. So this diagram here, I'm just simulating. Hey, request. get, https, google. com. We're pretty most of us are are familiar with that. Important thing to mention is that the kernel is the one that facilitates the actual transfer of the network packets and hands them back to the program. Python is not actually doing everything here. Python is at the low level communicating with this kernel And then the kernel's job is to go fetch those packets and bring it back. And yes, I know we're not.
Speaker 1: I didn't talk about DNS. We're not going there today. And it's a very simple thing that the kernel does. It's described in that diagram. Um so we're not going that deep. But the nice thing about that is Python C Lua, Rust, all these languages, they don't all have to re-implement the TCP networking stack and deal with getting these bytes from Google and grabbing them and reorganizing them in the correct order. No, the kernel does that for most programming languages and then hands them back the bytes and packets. And then the low-level Python code constructs that into a response for you. And as a Python programmer, which I love being one, I get this nice response object.
Speaker 1: I don't have to think about any of that because the kernel did that for me. And something I want to emphasize in this talk is there are certain things that the kernel does for your program, and that's really important to know about when we get into free threading. That the operating system is responsible for this, not Python itself. Okay, so skipping through that diagram, your program makes the request. The kernel facilitates fulfilling the request. Notice I didn't say the kernel fulfills the request because There are some people here that are quite smart that might be like, well, there are some exceptions. So I'll say it facilitates you getting you the data back. Same thing with files, you request an open file, the kernel gets that for you. You're not, Python's not reaching into your hard drive and grabbing the sector and saying, oh, there's the file.
Speaker 1: And the other thing the kernel does is prevents uh uh oh might need to restart my computer That is called a kernel panic. And because there is one piece of code running all your programs, if it crashes, you're in trouble. And that's what happens. I don't know if anybody else here has ever gotten a blue screen of death, but I do program sometimes in Windows. And it happens. So I asked an AI video generation service to illustrate the kernel for me. I did tell it to use a squirrel because I don't know if you all have studied squirrels. But I'll step away from that for a second. Squirrels are like woo, and then they go woof, and then you go do this.
Speaker 1: And that's kind of what the kernel does with your programs. Is Work comes to it to do, it does it, it hands it off. Work comes to it to do it, does it, it hands it off. And that's what's happening under the hood on your computer. And by the way, the kernel can do this with multiple CPU cores. it can assign work not just to one core, but to many. And those multiple cores can for the most part Run programs like at the same time. So the kernel can be delegating multiple tasks at the same time to multiple things, and multiple programs can be running at the same time. And one other thing to mention there is the kernel does switch between tasks. And I am not here to uh hypnotize you all today, but You can see
Speaker 1: what it would look like if the kernel was switching. Oh, that's interesting. That doesn't sync. Hmm, technology. There we go. There's my mouse. You can see what happens if the kernel starts switching programs really frequently. That's basically what it's doing. It's running one program for a little bit and then it's switching to another. And I'm not gonna sl drag that slider down all the way because Uh it's it's a it's a lot to take in. But eventually, actually to the human eye, it looks like that. I could tell you right now that the kernel is switching between the re the green program and the blue one 20,000 times a second. I don't know if it's actually that number, might be 200,000, might be 2,000. But it's switching between them so fast and it's running them that if you're like using your mouse or you're typing to you the computer feels instant.
Speaker 1: But really the computer's running, at least if you're on a Windows machine, didn't just check how many uh Windows Defender things are running right now But it's running hundreds of programs at the same time. And in order to do that and to still make it feel like the computer's fast, it's switching between them all the time. All right. So that's the kernel. That's how what we're going to talk about with the kernel. And some of those things I said are going to come back later. So keep that in mind. All right, process is pretty straightforward. It's just a copy of your program. And I'll do a quick illustration. Let's open up. We still have our local uh thing running. Oh, let's see how far I could zoom in here. A little hard to see.
Speaker 1: If I run Python here and I say x equals 10, and then I like, let's say I switch over to another terminal tab in the same folder, same Python. If I try and grab the value of x from this other Python program , well it's not defined. Have you ever thought like why is that the case? I'm running Python, I'm running the same program in two places with well, because the kernel's job is to isolate these programs. So it's not like running your source code, it's loading your source code into memory, then it's running it like a copy of your source source code there. And so what a process is, is just like a copy of your program running, is how I like to think about it.
Speaker 1: Alright, let's get those slides back up. Alright. So Nano Django has a nice way to hop into the shell. If you're running two shells at the same time, different terminals, those are two different processes. If you're running run server on two different ports, those are two different processes. And the kernel's job is to run those programs, and you know, you can exit out of your Python program and your kernel stops it. One way that you'll see processes in use a lot, let's get this mouse out there, is in GUInicorn, the standard Django deployment mechanism a lot of people use. At the top, here we have the kernel. The kernel makes a request, and GUICorn has this main process that grabs that and forwards it to
Speaker 1: one of multiple copies of your program running Which it used the kernel to create those copies. That copy processes that request, sends it back to GUnicorn, and sends it back to the kernel. And so that is how processes are used in Django. All right, so we got the kernel down, we got processes down, our crash course. It's coming along nicely. I'm over to one of my favorite parts though, we're gonna talk about threads, because we did say we'd talk about free threading. Two things I want you to keep in mind. One is that every process has a thread. So we'll talk about what a thread is in a second. It's better shown than laid out in words. And a process could have multiple threads. So what is a thread?
Speaker 1: Um, all right, this is showing up okay. So I have some source code here. I'll step through it a couple times, but Can you all see that arrow in the upper left? You may not be able to see the code well, but that arrow represents The kernel stepping through the source code one instruction at a time. And you wizards out there will tell me, well, it's not really stepping through that source code, it's stepping through You know, well, Python got compiled and then it's in the assembly and okay, yes, it's stepping through the assembly, but for the purposes of illustration. This is the operating system stepping through the c Python code one line at a time.
Speaker 1: I get down to some function, I call it, it steps into the function. If you've ever used a debugger, you've seen this. This arrow is the thread. The thread only moves forward. You can't move backwards even though I just did. When your program is being executed, it moves forward. It runs one instruction at a time. But there's only one arrow in a program if you run it. By default, there's one. But the kernel, and this is not Python. Python helps with this The kernel provides a way that this running thread
Speaker 1: can spawn copies of itself Within the same program. And I'm not gonna show you exactly what to call that is, but if I call this special function that tells the kernel, hey, can you make copies of me? I just need to tell the kernel where these copies start, like usually it's some function call, so we'll do that here. And ta -da! You'll see an orange and a green arrow. Colors may be a little rough. Now there are three threads in this program, and you're like, wait, no, I only see two. Well, the bottom one is still there. That's a a tough thing in threading for me to grasp is that all the original threat is still there. So now there's three
Speaker 1: Like threads, pointers in our program that the kernel can run all at the same time if it wants It could assign each of these arrows to a different CPU, and now we could have three cores running our this exact copy of the code that share that dictionary D at the top all at the same time. And I'm gonna take a big leap here, I'm gonna take a big risk. I'm gonna bring in the gill at the same time since you know it's a crash course, you might as well just introduce a bunch of things at once. I just threw a lock on that orange arrow, and some of you might feel this is backwards, and there's a couple ways you could explain it. What the gill does. Is
Speaker 1: it only lets me run one of these threads at a time. So only one arrow can ever be advancing at a time in the Python code. Now there are some exceptions, I'm not gonna talk about those. But this green arrow right now is the only one that's running. Oh now different threads running. The orange one's running, so the green one's blocked from moving. But Django's 20th birthday, what did Python get Django for the 20th birthday? Maybe it was the 19th. Ready for this? If you run Python in free-threaded mode, watch the screen, the locks gone. So happy birthday, Django
Speaker 1: from Python. Now, check it out. We just moved both cursors at the same time. We'll do some extra animations there for you to see it. So now we could have run at the one them running at the same time. One could run a little farther than the other. Now your operating system is running a lot of other programs as well. But that is my illustration of a thread for you. All right. And I mentioned I'd touch on there's standard libraries and a function in Python that helps you create threads. Not gonna go into that in detail. So why the Gill? Well, if two threads, we've got this dictionary here, if two threads try and assign to this new variable, this new key in the dictionary Z at the same time
Speaker 1: Suppose one's going to write the number 20 and one's going to write the str a string with 20 characters, happy 20th, Jenga. Well, let's pretend to be C wizards for a second. There are a lot of nuances here, but bear with me. Somewhere down under the hood, Python has to allocate memory for this number. It might only need one byte for that. For a 20-character string, it might need 20 bytes. And let's just pretend that the first thread Allocates one byte of memory. The second thread at the same time allocates 20. Uh-oh. Might have just overwritten it.
Speaker 1: And by the way, Malik won't do this, but Python might have pre-allocated this. It's probably safe, but this might be doom being done by the Python CE C level in the arena space. So thread where A writes its bytes, pretend there's some offset pointer, thread B writes its bytes, all of a sudden you've just written data outside of this memory region. Now I'm not saying this is exactly what happens, but this is what the GIL protects you against. It protects the low-level C internals from when they're extending, when the list is allocating more memory, extending itself. switching things around it protects that stuff from all sorts of uh-oh and so the gill
Speaker 1: made it a lot easier for you to You could write a C extension and assume, oh, I'm the only one manipulating this dictionary right now until I tell Python release the guilt. Okay, I'm safe to do all these manipulations. So it made it easier, it made it quicker to develop extensions, and I think it helped get Python to where it is today. And thanks to languages like Rust and advancements like their borrow checker , There was a conversation at PyCon 2025 this year. They're starting, Python developers are starting to look at, oh, there's ways we could do free threading but still have some form of memory safety. And so I think it's a great thing that free threading is happening now because there's been a lot of academic advancement. Alright. So There we go. OS crash course.
Speaker 1: Hopefully y'all are still with me. I don't know if that was 10 minutes, but it was pretty quick. Pretty quick. Okay, so is your Django code thread safe? Well, to answer that question, there's there's I like to think of it in three parts. Is Django thread safe? Are your third-party dependencies thread safe? And is your code thread safe? Is Django thread safe for the most part? Yes. I showed you with run server at the beginning, it's making a thread per request. And so Django's actually for a long time had to do a lot of work to make things thread safe. The main one, I'll say, was database connections. So a database connection in Django uses this special type of variable called a thread
Speaker 1: local, which can only be defined or read within a specific thread. So even if there are multiple threads, and here you can see uh in Django's database connections, when you request a connection to the database, it's actually aware of what thread you're in. And it uses a connection that belongs to your thread. So each database connection uses that, and that means two separate threads cannot access the same database connection. Now the lower level driver details, that depends on the driver. I'll talk touch on that in a little bit. But given the interaction with the DB is what, a third or fourth at least of what Django does, having thread-safe database connections is really important. And Django's been handling that for us for a long time.
Speaker 1: So I think we're pretty safe there. And um, by the way, uh yeah, it's I did want to mention Most Django requests, they go, you receive the request, you go execute some sequence of code, and then you send it back. Most things are doing things sequentially, one thing at a time. And if you're using async and joining all these other things, you probably already know if your thing's threat safe or not. And I know Andrew's here. I was at um DjangoCon in 2019 at the Sprints, and I kept overhearing Andrew and a number of other developers, they were working through the internals of the ASGI ref local. And figuring out like, well, what was the spawning thread that this database connection should belong to if I started in the event loop, went into a thread, then came back to the event loop, then went to another thread.
Speaker 1: And Andrew and many others did a ton of work to make sure that in Django, when you use it asynchronously or synchronously, if you call sync to async and you await that and you do this, Thank you, you all, for making it thread safe a long time ago because you're gonna see some numbers where you'll get a nice async benefit here. So the details of local could be a talk of its of its own, and I'm not uh qualified to give that talk. But you can actually see if you look at the source code for local, it has a lot written about threads and talking about thread locals and how this is actually an enhanced thread local because you know it does this, that, and the other. And so
Speaker 1: All right, enough about that. Are your third-party dependencies thread safe? My answer, I don't know. You got you gotta go find out because my third-party dependencies aren't the same as yours. But There's a whole Python team that has been working on helping third-party package maintainers publish a free-threaded build. They have a site out called pyfreethreading. github. io. If you Google anything related to Python free threading, you can find it. And there's tons of packages on there and they'll show you the various levels of support. And I personally love the site because I found like ten packages I didn't know about. They were just really cool as I was looking through. I won't go into the standard Postgres driver we all use, but there's actually work in progress right now.
Speaker 1: I know a lot of us use Postgres to make that officially thread safe. Um there's some areas internally there that it may not be quite ready for use yet, but you can follow that issue if you if you want to track that. Um and is your code thread safe? This is the last part. I'm gonna say potentially, and I'm even gonna say probably if you're just using pure Python and Django in standard ways. There's plenty of ways it might not be, but typically w if you're gonna write like Thread unsafe code. A lot of times you actually are pretty intentional about what you're doing and you would have almost already thought through that. Because Python still had threading for a long time. It just had the lock, so only one could run at a time. So
Speaker 1: I dug through the Django source code to find one example, and I found one of a thing that's not thread safe So Python has a cached property, but Django had it a lot long time before Python. And if you look at the source code of cached property, it doesn't use a threading lock. So I'm showing some made-up example here, but if you were to use a cache property and two threads were to access it at the same time, and I'm talking about using Django's cache property. Then you could actually get some unexpected side effects, like kind of like the ninja thing I was showing you, where this thing you think is gonna run once and then be cached actually runs twice because two threads access it at the same time. And this is a made-up contrived example, but I'm like opening some file, I'm incrementing some counter. Well, is your code okay with that happening twice?
Speaker 1: Because it could in free-threaded mode, unless you lock it. Um all right. We're touching on async plus threading. Uh async's a little more complicated, and let's see if you can see this diagram, but I'm going to talk about it with the standard Gunicorn approach of using a UV Corn worker. G Unicorn still has a main process and it still has these worker children processes But now all these worker processes run an event loop, not that much different from the kernel. But this is a Python event loop whose job is to handle all those awaits. That, oh, you await? Okay, the kernel, like, all right, I'll handle that and come back to you when I'm done. All right, go run another one that's running await. Go run another one that's running await.
Speaker 1: So this event loop runs your code, but then if it encounters a synchronous call with sync to async, it has to hand that work back to a thread. So here you have a lot of threads actually. In async Python, you are very coupled to threads because threads are there to help Django do work that's not like async native. And so that thread does work, passes it back to the loop, and maybe they do a few more dances, or a couple other threads get involved if you really have code like that. Um but eventually it gets back to G Unicorn, gets back to kernel, gets back to the to the client. Okay. I'm gonna skip that because I just went over that. All right, we're going to do the ASGI demo and wrap up with some conclusions. So my first one is ASGI
Speaker 1: G Unicorn plus UV corn with the GIL. So interestingly enough, uh didn't really distribute evenly to processes, and you can see some request times in the bottom left there, all over the map, um, which I find interesting. Uh the w the server waterfall is quite funny. Everything's happening at once because it entered all these event loops and then an await somewhere. And uh the next thing we ran here was let's do that, but let's do that in free threaded mode. You actually, if you look at the bottom left, Bottom left there, that's like request times all over the map. That's because of the Gill. If you look at it in free-threaded mode, you actually see very consistent performance, even with async.
Speaker 1: And to be clear, This is a CPU bound async endpoint that calls sync to async and awaits this really CPU intensive password hashing thing. Okay. So That one was pretty consistent and then Granion. Newer newer option. Granion's focused on ASGI, so this is where it claims to be quite performant. Um my initial experiment seemed to validate those claims. Um it did things actually a little bit faster. So If with ASGI with a CPU bound
Speaker 1: thread that had to run Free-threaded G Unicorn, which is the middle, was at least twice, if not maybe three times. Now I think two and a half times faster or so. And again, this is on a specific server configuration with a specific number of cores. This will vary based on your use case. Granian was a little bit faster. And this is the mean time, a little bit different per request. Granian was a little bit faster. The next benchmark is the most fascinating one to me, and I don't have a great explanation for you. G Unicorn with the Gill had much lower memory than G Unicorn in free-threaded mode. For some reason Running G Unicorn with Free Threaded spiked the memory, but Granian had lower memory than even GUnicorn with the GIL.
Speaker 1: So it was very interesting. And to conclude, should I use Free Threaded or not? Well, you you gotta first tell me, are you using sync, Django, or async? Because it's a very different answer depending on that. If you're using Sync Django today, uh free threading compared to say multi-processing, if you're CPU bound, it seems to have similar performance and lower memory. But it has more potential for bugs and issues, which these are not easy bugs. You probably don't want to deal with these And um if you're IO bound, uh similar performance, um, but it doesn't help you as much if you're I. O. bound. Um
Speaker 1: And in the async case, if you are CPU bound, I think it really helps the async case. We just showed that with the graphs. If you've got some CPU-intensive async work that you're running in threads. This is gonna be a benefit, but you still have more potential for bugs and issues, so be aware of that. And then I/O bound, uh a little bit less significant, and so Quick disclaimer, my demo app that I threw up, it shows an example of how you could start to look at this, but there are great profiling applications out there, Century New Relic. whatever you're using. I know the there's a number of newer other ones as well. And so I'd throw it up on a staging server before you do anything and look at latencies
Speaker 1: P90, P95, P99. And obviously, like it's got to work in the first place. And you gotta look at memory, you gotta look at all these things because they'll vary depending on your application. But last but not least, My opinion, 2025 Python 3. 14. If you're synchronous Django, I'd just stick with multi-processing for now unless I was in a super memory-constrained environment. I don't want to deal I don't want to be an early adopter on the bugs personally. And that's maybe just me being selfish. But um I would reevaluate next year. They're making this faster and they're Uh uh all everything's getting better. So on the ASGI side, if I had a CPU bound async app, I would very much actually consider
Speaker 1: Switching as soon as everything's compatible. And thank you, Ken, for this note. Ken mentioned Django templates can be very CPU-heavy. So if you're rendering some a bunch of templates, this could help a lot in those cases. And uh lastly from the um oh yes and then a note from Zaggs, heavy DB queries if you're loading a lot of things at once can also be CPU intensive. And then I. O. bound, I'd measure and probably just wait and reconsider. So that's all. I think there's So a little bit of time for questions, yeah.
Speaker 2: Yeah, we've got time for questions. Thanks, Michael. That was really good. Um any questions
Speaker 3: Thank you for that. This might be a little tangential, but I'm curious if you looked into subinterpreters and how that would compare with uh free threading or a combination of both.
Speaker 1: That's a great question. It's very funny to me that all these things were in the works for years and they all landed in the same Python release. What I would say on the subinterpreter side is the folks at Magic Stack that do uh it was previously called EdgeDB, I think it's called gel now. One of the core Python developers that wrote Async, he did an interesting demo at PyCon with Hive Workers. And this would deal more in the background worker space. But I didn't profile subinterpreters because the tooling is now, I think, just starting to emerge around them. Like it there wasn't a plug-and-play with My current Django app that was like dispatching requests to a subinterpreter. So I don't know yet, but I'd be very excited to try it when um I think some of the tooling is is more there.
Speaker 4: Hi, thank you for covering the subject. I think it's really cool. This is going to be a silly question maybe, but do you think Python is becoming more like Java as it gets free threading? And that makes changes to how the reference counting works and there's garbage collection and all this stuff. Like, are they are they converging? I was just wondering if you have any thoughts on that.
Speaker 1: That's a great question. I don't think so. Java had to go through a long period when they kind of released some of those locks of a lot of bugs, is what I understand. It seems like Python might be heading for the same thing, but um I think some of these proposals around regions um as it r relates to the concurrency direction it seems Python wants to go with this. I don't know. I I still think at the end of the day, Python is still a dynamic interpreted language. So even though some of the like syntactic constructs start might to start look similar. I I don't f I would not say they're converging or on the same track. And the last note there is really great language features end up in multiple languages. You see this all the time. And I'm not going to say async await was great, I don't have a strong opinion on that.
Speaker 1: But those landed in like three or four languages at the same time. And so You could see like when a useful construct comes up, another one was like trio and structured concurrency, you'll see them show up in multiple languages if it's a nice programming concept that can be multi-language.
Speaker 2: Do we have any more questions? Still got time Okay. Um Michael's still here, so if anyone wants to ask him anything uh in the hallways, please feel free to do that.
In the speaker’s local `runserver` demo, free-threaded Python used multiple CPU cores and finished about 40% faster than the GIL-enabled version. The exact improvement depends on the workload and machine.
Discussed at 4:19The GIL allows only one thread at a time to advance through Python code, helping protect CPython’s low-level memory operations. Free-threaded mode removes that lock so multiple threads can execute Python code simultaneously, but code and dependencies must then be safe for real concurrency.
Discussed at 23:00Django is largely thread-safe, including its use of thread-local database connections so separate threads do not share the same connection. However, third-party dependencies and application code still need to be checked, and some Django utilities—such as Django’s `cached_property`—can produce duplicate work when accessed concurrently.
Discussed at 26:49The speaker recommends checking the Python free-threading package-support site, py-free-threading.github.io, which lists packages and their levels of support. He specifically notes that work was still underway for the commonly used PostgreSQL driver.
Discussed at 29:59An ASGI event loop handles asynchronous work, but synchronous calls made through `sync_to_async` are handed to worker threads. Async Django therefore still depends on threads whenever it runs code that is not natively asynchronous.
Discussed at 32:18For synchronous Django, the speaker would generally stick with multiprocessing for now, except possibly in a severely memory-constrained environment, because free threading brings more difficult bugs even when performance is similar. For a CPU-bound async Django app, he would seriously consider switching once the application’s dependencies are compatible; for I/O-bound workloads, he recommends measuring first and probably waiting.
Discussed at 37:43Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026