Keynote: The Natural State of Computers by Amber Brown

This video features Amber Brown at DjangoCon US 2019 in San Diego, California, USA.

Keynote: The Natural State of Computers by Amber Brown
0:54:38
Published October 25, 2019
902 views

DjangoCon 2019 - Keynote: The Natural State of Computers by Amber Brown

TBD

This talk was presented at: https://2019.djangocon.us/talk/keynote-amber-brown/

LINKS:
Follow Amber Brown ๐Ÿ‘‡
On Twitter: https://twitter.com/hawkieowl

Follow DjangCon US ๐Ÿ‘‡
https://twitter.com/djangocon

Follow DEFNA ๐Ÿ‘‡
https://twitter.com/defnado
https://www.defna.org/

Intro music: "This Is How We Quirk It" by Avocado Junkie.
Video production by Confreaks TV.
Captions by White Coat Captioning.

Summary

Computers spend far more time waiting on memory, storage, and networks than executing Python, while CPU improvements increasingly come from adding cores and distributing work rather than making a single core much faster. Amber Brown argues that asynchronous I/O lets one thread keep working while external operations are in progress, avoiding the thread limits and overhead of traditional blocking servers; this is why Pythonโ€™s async support, ASGI, and asynchronous Django matter. She explains event loops, non-blocking sockets, coroutines, and `await`, then compares threaded and asynchronous web servers: async has overhead and may be slower on low-latency local requests, but handles high-latency workloads better because it does not exhaust a fixed pool of waiting threads. The remaining limits are CPU-bound work and single-core performance, which can be addressed by using multiple processes or delegated worker pools.

Key takeaways

  • Asynchronous I/O avoids wasting a thread while waiting for network or other external data, allowing one event loop to manage many in-flight operations.
  • Pythonโ€™s GIL prevents ordinary Python code from scaling across multiple threads in one process, but it does not prevent asynchronous I/O from improving concurrency.
  • Modern CPUs provide more cores and bandwidth but not proportionally faster single-core execution, while extra layers and distance increase memory and I/O latency.
  • Pythonโ€™s `async def` and `await` hide callback and future machinery while returning control to the event loop during waits.
  • Async servers can lose to threaded servers for near-zero-latency local work, but they scale better when requests spend substantial time waiting on remote services.
  • CPU-heavy work still limits an event loop; multiple processes, sharding, or worker pools can spread that work across cores.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction and ASGI Amber Brown frames the talk as a follow-up to earlier Django Channels work and introduces the emergence of ASGI and asynchronous Django.
  2. 4:13 Asynchronous I/O Fundamentals The talk defines asynchronous I/O, explains blocking, and discusses Pythonโ€™s Global Interpreter Lock and non-blocking techniques.
  3. 10:00 Async Python in the Mainstream Brown reviews Pythonโ€™s evolution from generator-based coroutines to async/await and explains why Djangoโ€™s adoption of async matters.
  4. 12:15 The Limits of CPU Performance The discussion turns to plateauing single-core performance, security-related slowdowns, power efficiency, and the future of processor architectures.
  5. 16:56 Multicore and Chiplet Architectures Brown examines AMDโ€™s chiplet-based CPUs, cache organization, I/O dies, and the industry shift toward more cores and processors.
  6. 20:55 Asynchronous Computer Hardware The talk explains how independent clock speeds, USB polling, PCI Express, and processor interconnects depend on asynchronous communication.
  7. 24:36 NUMA, Latency, and Cache Coherency Brown covers non-uniform memory access, architectural complexity, cache locality, and the enormous latency differences between memory, storage, and networks.
  8. 32:12 Event Loops and Non-Blocking Sockets The talk moves from unavoidable waiting to non-blocking sockets, readiness APIs such as select, epoll, and kqueue, and event-loop design.
  9. 34:29 Async Python Web Applications Brown compares conventional Flask-style code with Twisted-based async code and explains how await hides the underlying callback machinery.
  10. 36:50 Concurrency Benchmarks and Scaling Benchmark results show when async systems outperform threaded servers, while Brown discusses thread exhaustion, single-core limits, sharding, and process-based scaling.

Transcript

9,096 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:15

Hello, um I am here to talk about the natural state of computers. And even though it ends up being something like this most of the time. I'm going to go into a you know a bit more about how things work before it gets to this point. So I am Amber Hawkey Brown. Uh my Twitter account is at Hockey Al. Uh I'm from uh Melbourne, Australia, so it's been quite a long trip uh to get here Look, when you live in Melbourne, it's hard to get a picture of Melbourne. So I had to make do. This is from when I went to Pie Cascades. Seattle, Melbourne, the same thing. Good copy. Most people know me for my work on Twisted. So Twisted is an asynchronous networking framework, which will be relevant to my talk.

1:05

But it's been around for quite a long time and I am it's uh release manager and one of its uh uh longest standing contributors. So this is sort of a bookend to a talk given nearly five years ago at this point. at a conference called Django Under the Hood, of which the Deep Dives Day is sort of inspired by. God, I look young. Yeah, still do. I got carded the other day. Anyway, uh so this was given in twenty fifteen and uh around that time uh we had Uh Django Channels was nearing 1. 0, was 1. 0, it was in development. So um the early version of Django Channels was uh

1:53

more oriented around keeping Django sort of synchronous as it was. Uh but having a sort of asynchronous shell around it so that you could do asynchronous things without having to care so much about your actual Django being natively asynchronous. So I've got a couple of slides from that talk, and at that point I was I was fairly convinced that the original channels would kind of That would work out because it seemed relatively low effort, low impact, and if I know anything from software that Doing too much at one time usually ends up in failure. But I sort of was thinking about a sort of a an alternate situation from what I was assuming was the foregone conclusion.

2:41

And that conclusion was that we needed to replace WISGI. Because WISGI is, at its core, a s a synchronous-y style interface. You don't have an asynchronous event loop or anything around it, and you generally run in a thread, and it just didn't work for that sort of asynchronous feature. You could do responses asynchronously. But it didn't work for asynchronous protocols as a whole, like WebSockets. Now that has come to fruition, which is ASGI 3. 0. It's stable, it's well defined, and it's got multiple servers. So I think we're in a world where that whiskey too sort of does exist. I also talked about how maybe in this world where we have an asynchronous whiskey

3:28

that Django might be able to thaw both asynchronous and synchronous views. And that could be pretty interesting. And I was considering at that point that maybe a native asynchronous method would be the long-term uh way forward. But it wo would be rather invasive. And later on Andrew will be talking about how uh how that project has been less invasive and not required hopefully any broken eggs at all. Maybe some fried ones, but no broken ones. So think of this talk as a bookend to that 2015 talk. Because 2015 we didn't have what we have now. Now we have in development an asynchronous capable Django. ASGI

4:13

is developed and stable, and we have multiple implementations of ASGI clients and ASGI servers But let's start from the start and talk about why asynchronous stuff is a big deal. Now when I say async uh through this talk, I mean asynchronous input and output. It refers to specific programming techniques that minimize the amount of time that your code blocks waiting on external data. Blocking can be quite bad. When responsiveness is key, such as in graphical interfaces, blocking on waiting on external data or blocking on heavy processing can turn your application into a stuttery, slow mess. Usually the solution is to move any sort of processing into an alternate thread, leaving your main UI thread

5:02

uh to be very snappy, very responsive. And just delegating anything that would take any significant amount of time to worker uh threads or processes. Unfortunately in Python it isn't that easy. We simply can't make a thread on put on another core to double the amount of work we can do because of what it's called the global interpreter log. The global interpreter lock is both an essential part of Python and makes it very easy to program in Python and have some threads occasionally without having to learn entirely about how to write thread-safe code, but is also its Achilles heel. It allows Python code to execute on only one thread in that process at any one time, meaning that we don't have to worry about one bit of Python changing a list or a dictionary from underneath us.

5:50

which can happen with uh asynchronous things in other languages. You usually have to have a lock and make sure that things, you know, all know that things underneath it are changing. In Python, we don't have to do that. The Gill is a very rough lock, as the global in the name might hint. It means that we can't simply say that this code over here doesn't need the guarantees it provides, or that it safely locks. Because if we were to turn it off, code that we depend on and call would likely break horribly. Because when you're writing Python, you're not only writing your own code, but often losing lots and lots of other libraries that maybe don't have those sort of I'm okay with being multi-threaded. And a lot of Python code is not multi-thread safe. So for now we're stuck with it.

6:37

So if we can effectively only run Python on one thread at any given time, how can we best make the use of that? Well, if we only have one thread and do all the work on that, we don't have to pay all the costs of threading, like thread switching and locking and memory overhead. But if we were to block when waiting for external data, we'd only be able to request uh to process one request or one thing at once. But instead we can use non-blocking techniques for our I. O. and process as many requests as we have the CPU time for. So the solution to blocking and only having one thread is rather than using uh rather than waiting for the data to arrive, we use the operating system's non-blocking I.

7:23

O. functionality. It's not instant since we have to shove the data off to the operating system, but it blocks as little as possible. When we want to receive data, we just have to check for it instead of stopping the world and waiting for everything to arrive at once. Async isn't new. And because asynchronous I. O. underpins basically every high-performance networking system there is, operating system support is nearly universal. Of particular are the ePole and KQ APIs on Linux and BSDs and I. O. completion ports on Windows. And it's not new in Python either, because Twisted and Tornado have been using these interfaces for many, many years. And G -events and similar things like that being alternatives for almost as long.

8:09

Twisted and Tornado are quite similar and built on the concept of an event loop or a reactor, which we we'll get into. Well GEvent was a bit different and used a sort of greenlet model rather than an explicit event loop. Greenlets are similar to threads and have a lot of the same downsides, but are managed in user space instead by the operating system. Which gives uh your program more control. All these options are very mature and in wide use by sections of the community that need them. Unfortunately, it is a little bit hard to do asynchronous code in Python due to its lack of first-class language support. Back in a you know the Python 2 days, you could implement coroutines, which uh make it you uh make it a lot easier. You could implement it using generators, but they're slow and give messy tracebacks since generators work on uh exceptions

9:00

to finish uh finish them and they're very easy to misuse. One of my coworkers has a post-it note on his monitor that says, have you remembered to yield? Since when using these old sort of hacked in coroutines. Forgetting to yield on something, Python had no way of telling you. It didn't know that you were supposed to do that. Fortunately, everything's a bit better now. We have real coroutines, real-ish, and uh you can write them using async dev. And then inside that, instead of yielding, use await. To the user, the flow looks like you're blocking while waiting for the thing you will wait it to finish. But the await keyword actually gives control back to uh

9:46

back to Python and the asynchronous system. Meaning that you're not actually blocking. You your code just seems like it is. So it's much easier to follow if you're, you know, used to regular synchronous techniques. It's much easier to understand than your callback method of doing it, and it's a very familiar sort of API. Of course, there's new kids on the block as well. AsyncO was designed as a common kernel for all the asynchronous systems in Python, but has kind of become a thing on its of its own. Frameworks like Trio turn the conventional systems on its head, using the new native coroutines to provide a bit of a different API for your non-blocking I. O. But I'm here to talk about Django as well. Django 3. 0 will come with native asynchronous support for quite a few of its APIs.

10:35

Meaning that asynchronous I. O. is going to be something that innumerable Python developers will now get to take advantage of on a day-to-day basis. This is really unspeakably huge as it brings async truly into the Python mainstream. The effect of async IO and Twisted is going to be relatively small compared to uh frameworks such as Django adopting these uh systems. But if Twisted's been around for nearly 20 years and Async. io has been around for five, why is it only now that something like Django is adopting its ideas? The benefits of using asynchronous I. O. especially in web development, have been clear for a long time, but we're sort of uh approaching a tipping point. It seems now that async is no longer something that you can just opt out of supporting.

11:24

After Node. js came onto the scene, the tide sort of started turning. Node. js in its uh early days was very similar to Python at the time, where it had little language support, uh native language support for doing things asynchronous things, and it was bound by single-thread limits. But despite that, people were able to do lots of new and exciting things with it. We all know now that async in JavaScript is more or less as good as you can get in an interpretive language. Interpretive rather. Funnily enough, it was actually JavaScript that adopted Twisted's ideas first. What we know is Promises A plus actually started out as a port of Twisted's deferred to JavaScript. Which was then subsequently iterated on through different frameworks and standardized to the form we know today. Now, upon seeing sort of how good async could be, everyone sort of wanted in on it, wanted this native support.

12:15

Languages that didn't have a good story at the time, like Python, ended up losing users to ones that did, like Go. So first class support for asynchronous code seems to be a must. Last month, the Rust 's long-coming async await support landed. And I'm very much looking forward to see where that goes. But what is this requirement for an eSync story driven by? Is it just people wanting this fancy new feature, or is there some technical reason why it's so effective? I think it's that computers are not getting faster. We're having to be more efficient with what we have. And for a lot of networking-based workloads, which which a lot of web development sort of goes around

13:03

focuses on, especially when you're talking to web clients and databases and all of that, most of the time is not actually in Python, but in networking. And when you're doing that, asynchronous sign is far more efficient than that of a traditional blocking regime. So you might think computers are getting faster. Very much so. Why would we buy these new ones if they weren't getting faster? This is the history of you know mid-range Intel CPUs for the past ten years, and as you can see, the geek uh the passmark numbers are going up. Faster, case closed, right? But benchmarks don't tell the whole story. Those jumps in performance are because those CPUs added more cores.

13:50

If we divide that performance by the number of cores and the clock speed, We can see that there hasn't been a significant increase on the whole since the Core i5660 launched in 2009. In fact, performance has even dipped a bit for the past couple generations. Why is this? Spectra and Meltdown has sort of ruined everything. They wound back nearly every relative performance improvement that Intel CPUs have had since the call uh uh core series' uh inception. And even the 10th generation Intel CPUs due to launch at Christmas won't really improve it that much. All these improvements that we've been getting are based on things that are insecure and we're now having to roll them back. Of course, we have improved technology to the point where there is a point to new CPUs.

14:37

They use less power. And that means that even if our CPUs are the same performance clock for clock, we can still run them at a higher clock speed to a limit, add more cores, and stay within the same power envelope. That's a massive boon for consumers, as even a low-power CPU can turbo boost to four gigahertz and higher for short times. But it doesn't really help those of us that deploy to servers or do enough work that the frequency boosts eventually have to stop because you're uh causing too much heat. And I don't think that this is just the past ten years. I don't think that the next ten years is going to get as fast as computers either. ARM is looming on the horizon, ready to take over the server market. But that's because of price and power efficiency, not speed.

15:25

You can fit more theoretical performance per square foot with an ARM cluster, but each one of those individual CPUs is slower than the equivalent x86 core, and might well always be. Other contenders like RISC-V are promising but yet to be proven. The first few letters in RISC stand for restricted instruction set, in contrast to the complex instruction set of CISC processors like x86. Intel 's early solutions to performance problems was to introduce instructions that did more per single instruction, therefore getting more performance out of less instructions per clock cycle that you get in a Cis regime This let them dominate the computing space for the next few decades, but the complexity of CISC processors now makes the relative simplicity of platforms like RISC-V

16:10

attractive. Not for performance or really for security, but for the ability for a human to maybe understand what the hell is going on inside of it. It is unfortunate though that we are ultimately hitting the limitations of our control of physics. The smaller our transistors, the harder they are to make, the lower the yields, and the more difficult the next iteration. We have taken every easy option available in the hardware, so it's time to start taking the harder ones. So what are the kind of processes and platforms we as Python developers or developers at large will have to operate on in the future? We can look at the present for what the future holds. This is the Ryzen 5 3600, the CPU that's in my desktop at home. It's actually a picture.

16:56

It's a consumer CPU, but it's very similar to the high-end and server processors that AMD are releasing. Just use uh they're using the same core architecture and even the same silicon, just arranged differently on larger chips. This CPU with six cores here costs 200 American dollars, while the Epic 7742 you might find in a new uh render farm server costs 7. 5000 US dollars. But that has 64 cores and 128 threads. They have a lot of products in between, but it's worth looking at what we'll do with those ones in between. And it's likely that you'll find things similar to this or working on the same architecture in render farms and general purpose servers over the coming months.

17:42

But again, it's not about any one CPU or any any one series of CPUs. It's a small-scale replica of where the industry is going. Multiple and more CPUs and cores on a single processor. And all of the interestingness that that implies. Intel has announced a similar topology in their future server CPUs, having multiple chips. Inside the same processor. And it is logically similar to the existing multi-CPU boards you can find in servers. Now, if we look under the hood, uh this CPU has an IO die, which is on the left, and a chiplet, which has six cores on it. You can see underneath the text there's a bit of a room for a second one, and that sort of is the core of this architecture, is that you have the IO

18:30

die, and then you can add diplets to get the number of calls you want. On each one of those core die triplet dies, there are two core complexes. You can see them vertically. Each core complex, or CCX, has 16 megabytes of L3 cache. And 512 kilobytes of L2 cache for each of its four cores. Since the phase 600 only has six cores, two of the cores on this die are fused off and inoperative. MD has built their entire CPU line, server, and consumer around these eight core chiplets and use the ones that have faulty cores but otherwise work in lower spec products, similar to the uh Athlon X3s. That were about 20 years ago now. They were four core chips that had one faulty core, so they sold it off and made a three-core chip.

19:19

Now they're using these eight core triplets in everything from their consumer CPUs to their server CPUs. In something like the Epic 7742 I mentioned before, it would have been a single of those IO dyes and eight of those fully operational chiplets to provide the 64 cores. As we can see with this little diagram, the core complex has a chunk of L3 cache and two cores. They're associated with an L2 cache, uh mirrored on each side. So you can see the two cores in the L2 and the L3, and it's the same on the other side. The Iodie is sort of marketed by Intel as controlling what they call the infinity fabric, which is basically just a marketing term to refer to the high-speed connection between the Iodie and the chiplets. Which scales up depending on how many triplets you have.

20:07

On this single triplet CPU it doesn't really have a lot of benefits, but it's important to think about architecturally. Now the IODI is responsible for connecting to your RAM, as well as being what physically hosts the PCI Express lanes. It also has some high-speed USB and that sort of thing. thing. That's the responsibility of this die and not the chiplet cores. The chiplet just has CPU, while this I. O. die has the other peripheral things that you would expect a CPU to host. Some of those PCI lanes then go to the motherboard chipset, which then hosts things like SATA and more USB ports. And all of these parts have to communicate asynchronously, and it's a hard requirement that they have to. Asynchronous communication between your motherboard and your CPU

20:55

is core to allowing different parts of your computer to get faster without requiring the others to act in lock lockstep. Get some water. Delicious. Nowadays, everything on your computer will have a different clock speed. PCI Express has a different clock speed than your CPU and your RAM. This wasn't always the case, as the older frontside bus design of of 90s and early 2000s computers had a frontside bus clock speed and then a CPU clock speed multiplier. Your frontside bus clock speed often had to be the same as your RAM clock speed, which would introduce a bottleneck, especially since your frontside bus clock speed on some systems

21:43

would also dictate the clock speed that your PCI devices had to operate. So this one clock speed not only made how fast your PCI was and your RAM was, but also how fast your CPU was. Bottom computers support all of these having different clocks. And to have different clocks, you need to run them and communicate asynchronously. In fact, the front side bus on computers no longer exists, and the CPU will often have its own memory controller. And will interface with things like PCI Express and et cetera directly. On the newer CPU architectures like Zen2 of the Ryzen 3600, even the clock speed of the Infinity Fabric interconnects. between the iodye and the chiplets can be independent of the iodie and the chiplets, meaning that you can overclock those chiplets without overclocking the iodye, or vice versa.

22:33

Who remembers these? I don't. I lie a little bit Nah, I was an Amiga kid. Anyway, PS2 keyboards of mice, the the real, real-world ones, operated on an interrupt system. When they had something to report, like a mouse movement or a key press, they would send a CPU interrupt to tell the CPU that something had happened. It kind of is what it sounds like. It interrupts the entire CPU to get its message through. This, as you can guess, kind of sucks if you're trying to do other things with your CPU. So we eventually replaced it with asynchronous systems such as USB. USB is instead a polling -based interface, where it's up to the host computer to query things from a device to make things a lot simpler as well.

23:18

For certain devices like mice and keyboards, the instant response of interrupts was advantageous. But with modern gaming mice polling at like a thousand hertz, this is no longer really a problem. And it means that when we have multiple cores, we're not shutting down the world to tell you that your mouse moved. Polling doesn't quite work with everything though, especially higher performance uh platforms. Things like PC iExpress use a bidirectional asynchronous communication system, which is almost comparable to something like Ethernet. You send packets of data down to the PCIe device and it sends packets of data back, like it like it were, speaking Ethernet. Getting these messages doesn't halt the CPU and uh with like these newer things with the ITI, it's actually quite far away from the CPU cores.

24:06

The PCI data is then uh funneled over the interconnect to the CPUs, which again is running, can run at a different clock speed than both the CPU. And the PCI Express bus. This gets relatively important when you look at, for example, the new PCI Express V4, which is even higher speed Because if you were limited by how fast your CPU was, we wouldn't be able to get these improvements in hardware without having to have entirely new generations of CPUs. that run much faster. Because we can't get much faster with the CPUs, we have to, you know, untie them. Of course, the IODI being central and separate to the CBU cores sidesteps a significant problem, and there's NUMA. NUMA stands for non-uniform memory access. And describes an architecture where you have multiple CPUs that don't all share the same memory.

24:56

You can have a large common memory bank, but that can increase latency as the memory gets physically further away from the CPUs on the motherboard And the chip, and more CPUs are using a limited common bus. Instead, when you have lots and lots and lots of cores, a Numa system will have some CPUs having local memory, which is very close and high performance. And an ability to ask other CPUs to pass through memory that it controls. Now, the previous generation of AMD's high-performance systems implemented Numa, and it was a real brain scrambler for traditional software. You had close memory and far memory, and a lot of software that we write doesn't interact well with that. In Python, there's no way of saying, no, no, no, keep things over here and avoid calling out

25:45

over there. And this was the case for games and that sort of thing. This got better because uh Windows and Linux and that decided uh decided to keep things in the memory that they were dealing with, but early on it would just put them anywhere and you'd end up paying these pneuma costs without any real benefit. Now, that's not the case with these newer CPUs. But if you have multiple CPU sockets per machine, you still have those NUMA characteristics. Because each CPU will have its own memory banks, even if all the cores share the same memory bank on its CPU. So what does this leave us with? We have more cores, which is great. Everyone loves more cores. We can access more memory, especially on dual socket machines, because of improvements over the past 10 years. And high-powered interconnects means that we have more bandwidth to everything, such as PCIe or RAM or whatever, or between different chips.

26:38

But these performance improvements have an architectural cost. We've got lower single-core performance, especially on alternative platforms like ARM. We've increased latency because things like the IO die are simply more things between us and the things we're accessing. And even if it allows us to access more at a higher bandwidth, we're still gonna have to wait longer. And it adds a bunch more complexity at every level, especially when you factor in things like cache coherency. So let's look at one of those things, latency, because that factors into networking. It's important when you think about asynchronous I. O. but it's also important to recognize what effects latency has on the different things we might want to do on the c on local to the computer. So 3. 6 GHz CPU, like my one, does 3.

27:26

6 billion cycles per second, each taking 0. 7 nanoseconds each. Now the L3 cache on that chip uh takes about 40 clock cycles to communicate with, which is about 11 nanoseconds. The DDR42666, which is fairly decent RAM, will take 55 uh clock cycles to access, and that's Pretty good as well. Now the fastest SSD you can buy, and we're talking tens of thousands of dollars here, will take 220,000 clock cycles, or 60,000 nanoseconds. So it's quite a big jump from cache and RAM to SSD. Now my local router to my local ISP takes about 12 milliseconds or 12 million nanoseconds.

28:14

It's about 45 million clock cycles. Now if we do Australia to Los Angeles, which Melbourne to Los Angeles, which was the flight I took, for a person that took 14 hours. For a packet, it takes 350 milliseconds, and I can tell you which one I would prefer. But that's like 1. 3 billion clock cycles. That's you know getting quite large. Think of it as like the cache is like your desk, you grab the paper. The DDR is the filing cabinet, maybe in the next room, depending on the speed of your RAM. The SSD is, you know, an in-state parcel coming in. And sending something to Los Angeles is quite literally just flying to Pluto. It's it's like most of a human lifetime in comparison.

28:59

So how do we make the best of this kind of latency? Well, when I said that the Ryzen 3600 had six cores, it also has 12 threads. And that means that supports simultaneous multi-threading. Now, simultaneous multi-threading is the practice of having two virtual cores or two or more virtual cores per real core. with the assumption that one of those virtual cores will be waiting on RAM or disk to do anything, meaning that the other can actually use the processing parts of the CPU while the other one is waiting. It's not something you can really control, but it's worth knowing that some workloads operate better with it and some don't. In the context of Python, because we're accessing things like the RAM a lot and the DIT ,

29:48

You can get about 150% of the performance of a single core with this because you're simply slotting things in when they're ready and utilizing that actual processing part of the CPU more effectively. It's also important to maintain cache coherency. With mod CPUs having upwards of 16 megabytes of L3 cache per set of cores, keeping that cache valid and close is extremely important. By staying on the same core or the same core complex, it will stay coherent and fresh and close. And that means that with such a large amount of cache, you can store significant parts of executables in your cache without having to load it from disk again. Things like sending CPU affinity, keeping processors local to their cache, therefore

30:34

can become extremely important for getting the most out of those CPUs. So the best thing to do in this sort of scheme is having uh every available working core have only one thread running on it, staying as active as possible. But sometimes waiting is unavoidable. Those fetches from RAM or disks will block your application no matter what you do. And it's not like we can avoid fetching from RAM, we're avoid fetching from a disk. Sometimes there's things that we have to do. But the point of asynchronous I. O. is avoiding it where you can. And networking is a great place to where you can avoid it. Preemptive multitasking can be a bit of a mind-killer here as well, as the operating system will see that potentially you're waiting on some RAM

31:20

or CPU or hell even networking. And it will attempt to go, oh well, you don't need a CPU, so I'll put something else there. Setting CPU affinity means that you won't get moved to other cores, but the operating system might put something there that overwrites all your precious caches while you're waiting for a disk read. So there's things you can do in your operating system to avoid that, but basically just reducing the number of things that can be scheduled. Uh there's having lighterweight servers that run just what they need to is also important. But again, waiting is unavoidable. Fetches from RAM will block you no matter what you can do. But the point is sitting not sitting around for longer than we need to. And how do we do that? By using asynchronous networking techniques, like I've been going on about, we can avoid the blocking to be just those things that we can't avoid.

32:12

But keeping the networking, which again takes many, many, many, many times longer, making that not take, uh not having that not block. At the last level, we want non-blocking sockets. These will not block when we try and do things with them, like reading and writing. We'll instead rely on us checking on their readiness state before we try and do it. If the socket has no data, then it'll raise a wood block error, telling us to come back later. As well as if we try and write too much data to it and the operating system's write buffers become full, it'll tell, nope, I can't accept any more data without blocking. So we have the problem where we can't always talk to these sockets. So how do we figure out when we can? Now

32:57

select is the standard Unix API for working with these non-blocking sockets. There's improvements to the same basic formula in KQ and EPOL, which are the things that you know actually use, but the main uses, the main core of it is the same. You give it the sockets you want to know if you can use, and it'll tell you which ones can be read from, ridden written to more, or have errors, such as being unexpectedly closed. With this basic API, we can then implement what's called an event loop. The parts of our application that then use these non-blocking sockets register up on this event loop, and the event loop will then notify them when the sockets are ready. Turning this more into sort of a, hey, tell me when this is ready. Okay, here's the event which is data.

33:43

The real final analog is leaving the microwave to do its thing until it beeps, rather than standing there and waiting for it yourself. I mean standing there is simpler. You don't have to go anywhere. You can just look at it. You don't have to do anything. You don't have to worry about the microwave beeping while you're in the middle of some other task. But you are being inefficient with your time. Being natively async lets us work within the greater asynchronous system there's a computer and avoid wasting time. We are taking advantage of its natural state of asynchronosity. So let's see see how to do in Python. Go have a drink first. So let's look at conventional web

34:29

app. And I'm sorry, but this is floss because I couldn't fit a Django app on one screen. I love you, but you have a lot of code. But this hopefully should get through my point. So we have the Flask app and we have a route and we have a main and we do request. get to fetch something from say localhost, which is why I've been using for my benchmark. So I've got a local web server running, that's uh enginex. Um and then we return the content from it. Now, how does this look in an asynchronous web app? So Kline is sort of a uh implementation of Flask-like API, but using twisted and asynchronous things.

35:14

So as you can see, the code is very similar. We have track, which is like requests but twisted. We're great at naming. Um and instead of just calling it, we have to await it because it does some sort of networking. Now I'm using def async and await there, which is syntactic sugar for callbacks. If we were to look at the raw uh deferred callbacking way of doing it, this is what it would look like. It's you know not much more complex here, but with larger code bases it does make a lot of sense. And using a callback system with like loops can be quite difficult. So async def and await makes things quite easy. Now, the core butt of async code in Python is that it'll end up returning something like a promise or a deferred.

36:04

It's usually some sort of shell that says, hey, in Python, I kind of have to return something, but the value's not here yet. So here's a box that will have the value at some later date. Now in a coroutine, uh you when you await it, you actually await that promise or deferred, which will then suspend the coroutine at a callback. or such onto the refer uh return deferred or future and then resume it when that triggers. So all of that callbacking sort of happens invisibly to you. Now, can't see it on here, so I'm gonna Okay. So this is a chart of our concurrency. So on the left is requests handled per second.

36:50

And on the bottom is the number of concurrent requests I'm throwing at the server at once. So this is on my Resin 3600, so it's quite beefy, so um I don't expect to run into any problems where I'm hitting CPU lib. quite yet. So as you can see, being asynchronous actually is not doing well. It's actually slower. Now why is that the case? Now, when we have low latency, such as communicating with localhost for this benchmark, we block for less time. And working on a local machine like this effectively has zero latency. meaning that the threaded solution will not actually block much more than the asynchronous one. And because there's not much blocking difference between them, the extra com uh computation handling the event loop and registering

37:37

you know, the sockets and all that sort of thing means that we actually spend more CPU time serving each request in the threaded model which bottlenecks us. But what happens if we change that zero milliseconds latency and introduce 350 milliseconds latency? Well, things end up being quite different. I actually had to add a second Gunicorn there with 12 threads and 12 workers instead of just the normal four. So in this case, the single-threaded twisted application handles many, many more requests per second. Now, the reason for this is threadful exhaustion. Because when you're running a Flask app in Gunacorn, you have a limited number of threads that you can run it in. And when you hit that thread limit where all of them are waiting for that licensee, you can't serve any more requests.

38:24

You have to wait for a thread to finish, and then you can put a new request in while twisted, and other asynchronous systems don't really care and can just keep adding adding them to the event loop. So you end up with oh god I've it doesn't show up on here, so I I need to remember what charts I'm looking at. So um the 95th percentile latency. Um you can see here that the licensee goes far up, so the 95th percentile is like the the uh worst five percent of requests. um is that Gunicorn will get really bad when you give it more requests and it starts exhausting it. Twisted will happen not far behind, but we've twisted on PyPy, which is a just in time inter uh

39:10

uh compiler interpreter for python, which lowers the uh CPU load, that it sort of levels out a bit because at that point we're running out of CPU. Um and that's what's causing Twisted's latency to spike. So we can see the concurrency limits here. With Gunicorn and other similar threaded systems, the hard concurrency limit is the number of possible threads you can run. With async, it's a bit different because it's your single-core performance. Once you start having more than one second of work per second, your event loop is no longer reactive and it falls behind and latencies go up. So we can add more threads, that always works.

39:57

But we can't always add more single core performance. So that sort of puts a hard limit on us. And in Python gel means no Python multithreading. So there's not really that many ways to make that better because we can't just start up a second event event loop in the same thread and have that handle. But there are methods such as sharding. Now you are limited by your single core performance, but there are tricks to make it so that you can have more processes which each have their own single core performance. limiting them. So multiple processes accepting incoming connections on the same socket works very well. So they're just round robins between the multiple services that are getting the requests, which means that then effectively you double your performance.

40:43

You can do this with some Linux APIs that will let you bind twice on the same socket, or you can have something like HAProxy on the front end that then directs it between multiple servers. You can also have a single process doing these accepts and then delegating it to subprocessors. And because those subprocessors are in different processes, the dual doesn't affect them. This works much better with not with this latency thing which where you do no work, but say you need to do a little bit of work and you're hitting your CPU performance limit based on that. By delegating, you can Level it out. But at the end, you're going to have a limit where you can't uh serve more raw TCP connections. You're going to hit that limit eventually. which is what we are seeing there where it was like several hundred.

41:29

And this is a desktop computer, and with faster servers, you might be able to handle thousands of requests a second before having to run into that. Now, these sort of things are workarounds and they don't solve every problem. But they do sort of turn some synchronous problems such as processing into asynchronous problems. If you have a subprocess worker pool or a distributed worker cluster, those things that you might otherwise do in your application, you can just put somewhere else. And this works very well if you have different kinds of workloads. So say your application is something that does some data forecasting and they send some JSON. And you get that JSON, and then you put it in the uh you decode it and you go, yep, this is a valid request.

42:16

Encode it again, put it on a work uh on a worker, send it off. The characteristics of your application that you're writing and of the worker don't have to be the same. Your worker could be running on something like Like PyPy, that's extremely fast and can serialize uh deserialize and serialize JSON very quickly, while your workers might be running on very beefy servers with lots and lots and lots of cores. uh running C Python with NumPy and C extensions and maybe GPUs. So by adopting this sort of thing for any significant amount of work you need to do that's heavy processing, you sort of allow yourself some level of scaling. But now I have a distributed system, you say. I've got a worker pool and all of that. And that's things that you kind of have to care about.

43:02

And it's too bad you already had one. Threads are a distributed system. And when we start up more threads, we kind of don't realize some of the implications that that means. If we're running something on a thread, it's like, okay, we don't have to worry about the uh the worker going away. But it is there are potentials where a thread can jam, where a thread can uh trigger an out-of-memory exception, that sort of thing So when you run into these limits where you have you know distributed system problems with threads, everything just falls down in a heap. Well if you do it properly, you have things like retries and locks. But I've been talking a lot about performance, and for me that's one of the main benefits.

43:48

I in my day job work on making networking systems that can scale to thousands and thousands of concurrent users. Not everyone has this problem. But it does, being asynchronous does give you some benefits apart from just being able to serve more users at once. Now I'm gonna say something controversial and say that server-side rendering is effectively dead. And if it's not, it's at least halfway to pining for for Jaws. We're instead shipping data to the client, not rendered HTML. If we are shipping rendered HTML, it's not that much Often, this is over things like WebSockets, which is bi-directional, or with server-center vents, leveraging longer-lived connections that the synchronous request-response cycle does not fit well with.

44:34

Being asynchronous means that we can not only communicate asynchronously with the things we're getting data from, but the things we're sending data to. And when you look at things like mobiles, That can be extremely important because when you have a mobile, setting up one TCP connection over a 4G network takes Several hundred milliseconds. If your web application requires setting up a TCP connection every time you open the page to load that HTML, that's going to be slower than if you had client-side rendering and a concurrent uh connection that just sent the data over web sockets. Things like HTTP2 sort of hack around this by letting you send you know blobs of HTML over one connection, like sending lots of pages over one connection.

45:19

But there are certain situations where sending the small potentially smaller data over the wire to those clients is potentially better. And we're not dealing with big data. Big data isn't really a thing. We're dealing with lots and lots of small data, very often. Sure, when, you know. it it's added up, it becomes a big data. But those problems are usually not our problems. And they're usually in the domain of the data warehouses and the people that care about like handling those big data, that big data. We're handling with the small data. We have to deal with being able to get it to where it becomes big data and doing things for results that come out. And doing that in an effective way.

46:04

So being asynchronous, we can send lots of little bits of data to various different systems and do that without blocking the web request. Message passing systems are great and simple to scale, relatively. Because talking to other systems, such as sending uh signals or sending messages to other parts of your uh distributed system. or writing to the database, for example, for statistics, because it no longer strictly blocks the response, we don't have to worry about making it technically slower. This gives a lot of potential for moving things we'd usually do in a request to a message queue, for example, to be picked up by an independently scaling system that's maybe more formant. But it also means that when we're doing requests, we can be like, oh, we need to update this counter.

46:54

That's no longer something that strictly means that you're going to make your request slower, because you can update that counter and serve the user's request at the same time. So Django. It's coming soon, thanks to the work of a very determined uh individuals in this room looking at one, but there are many others that are working on making this a reality. Yay! It's something that really excites me as someone that is not being a Django developer for quite a while. Because it means that now I can use Django for these applications that I would otherwise have to go and use Twisted or Asyncaio or something. for. And you know, all of these things aren't quite in there yet, but handling things like WebSockets and simultaneous database queries and l high-latency things like triggering webhooks is now something that I can just be like, okay, I'm gonna write this in Django.

47:51

I'm gonna have the nice things that Django gives me. I'm gonna have Django migrations, I'm gonna have Django middleware, I'm gonna have all this stuff that I like, and not have to sacrifice for it. Of course, this new sort of asynchronous Django is not built on WISCI, although it is backwards compatible with Whiskey. But for the native async stuff, you need to use ASCII. Hopefully I'm getting support in twisted soon so that you you know you can just set something up uh very simply. And this is kind of exciting because UVCorn and Daphne are Python. So we're no longer worrying about something like Nginx controlling our web servers. It's Python serving Python. Serving HTTP. There's just Python all the way down.

48:36

Which means that us as Python developers, we now control the whole stack. We get to do really neat things at the HTTP layer without going, well, that happens in Nginx. We can't improve that. We just have to wait for Nginx to fix it. We can go, well, we need we want to support HTTP3. Cool. Let's get some Python developers around and make that happen. It sort of gives our sort of brings the destiny for that sort of thing into our own hands instead of worrying about what everyone else is doing. And when you've got a server, you don't really have to just limit yourself to ASGI and nothing else. Because you're natively asynchronous, you can run other asynchronous things inside that process and have them all kind of work well together. One of the most useful things for debugging that I've ever used is Twisted's Twisted Contra Manhole.

49:22

It allows you to SSH or telnet into your running Python process and get an asynchronous REPL, like your teenage mutant Ninja Turtle debugging Manhattan. It's great and you shouldn't do it in production, but I do. Sometimes. You could also implement protocols like DNS and run that as another service inside the same process, having your async Django app serve the web interface and host the database. Now that seems sort of like, oh cool, you can do DNS, but when you look at the success of something like Pi Hole, which is a ad blocker that you can run on your Raspberry Pi or other things that blocks DNS uh dynamically. It actually kind of seems a bit more interesting because you could now do that sort of thing in Python.

50:09

When IoT means that you have all these little things around with all these protocols, It means that being natively asynchronous allows you to do a bit more. And the sky's the limit here. You're no longer constrained by what a whiskey out can or can't do. Because now you can do anything that asynchronous Python can. You can do asynchronous serial to an Arduino or Web to the world, DNS or serving SMTP or even XMPP. You you don't only have to be able to talk them, but you can now serve them from what is your Django app. And this is kind of important for Python in the sort of IoT space, because when you have your Django app You might want to talk to say Zigbee. And you can now sort of do that a little bit better because you're not worrying about the thread that goes away.

50:57

You can have your persistent communication and your app and then just have it all in one process. Maybe not the best solution, but hey, getting something working is better than something that doesn't work. So what kind of progress are we at? So DEP0009 has been accepted, which talks about the implementation of uh Django async. The SGI support has, as far as landed. Yep. Uh Django 3. 0 will ship with async views and async No? No? When did this change? Oh god. Okay , just ignore that. Um Django three point something.

51:43

Django three point zero on my heart. Just use Django Develop, it's fine. Um, so ignore the numbers here. Um I wrote this a month ago before this changed. So A Django will ship with async for using to wear. A Django after that point will ship with an async capable ORM, not an async native ORM, and potentially async templAU. And then the future is sort of upgrading the rest of the parts of Django to support, like uh asynchronously sending emails so it doesn't block your request, and having a natively asynchronous ORM that interacts in that event loop natively rather than putting it in a thread pool. But I want it now. Well, if you want now with Django, you maybe can just go

52:30

see Andrew and be like, hey, go. Got any async for me? I want some async. And then, you know, which I'm sure will be happening at the sprints. Yes. So at the sprints if you would like to get a taste of it in Django, go an OAndrew. And help him, for love of God. Love of God. So but if you want it now and you want it to play with asynchronous systems, there's a couple of options. So Twisted is the one that I'm biased toward So there's Cline, which is a WorkSoog uh-based API in Twisted. There's Trek, which is a request-based API, and TX Postgres, which gives you natively async Postgres. For async AO, there's libraries like AOHP and async PG, they'll let you do Postgres.

53:18

Trio, which is very interesting, has a H11, which is a HP one point one implementation and Shri OPG, which is async Postgres, and is something to perhaps look into and play around with just to sort of see what's possible, even if you don't you know use it in Anger. And at this website, this web zone, my web ring, um I've got some links to all of this various thing, various parts of things for you to to look at such as the DAP0009 and various announcements, as well as background posts from me and Glyph and others about asynchronous systems in general. Now it's been lovely being here and I'm loving you know being in the States again, sort of, as much as I can.

54:04

And This is my first Django con and it's been wonderful and thank you all for having me. And I hope you all have a wonderful day and that you all go see Andrew's talk where you see about how this is happening in Django in reality. And that's all. Thank you very much.

Questions this talk answers

Why is asynchronous I/O useful for Python web applications?

Pythonโ€™s global interpreter lock limits the usefulness of simply adding threads for concurrent work. Non-blocking I/O lets one thread handle many requests while they wait on networks or other external data, avoiding the overhead of blocking and thread switching.

Discussed at 6:37

Why is async support becoming important for Django and Python web development?

Web workloads spend much of their time communicating with clients, databases, and other systems rather than executing Python. As computers gain cores more readily than single-core speed, and as async becomes mainstream in other languages, asynchronous I/O is becoming difficult for web frameworks to ignore.

Discussed at 10:35

How do event loops and non-blocking sockets work in Python?

A non-blocking socket reports when it can be read from or written to instead of making the program wait. An event loop monitors those sockets and resumes the relevant application code when one is ready, allowing the program to work on other tasks in the meantime.

Discussed at 32:57

When is an asynchronous web server faster than a threaded one?

With almost no network latency, async can be slower because maintaining the event loop adds overhead without avoiding much waiting. With high latency, a threaded server can exhaust its finite pool of waiting threads, while an async server can continue adding requests to its event loop.

Discussed at 36:50

What limits the concurrency of an async Python server, and how can it scale?

An async server is primarily limited by single-core performance: once the event loop has more than roughly a second of work per second, it becomes unresponsive and latency rises. It can scale across cores by running multiple processes or event loops, using shared-socket acceptance, a proxy, or subprocess workers.

Discussed at 39:10

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon US