High-Availability Django by Frankie Dintino

This video features Frankie Dintino at DjangoCon US 2016 in Philadelphia, Pennsylvania, USA.

High-Availability Django by Frankie Dintino
0:20:54
Published August 24, 2016
1,874 views

High-Availability Django by Frankie Dintino

One year ago we completed a years-long project of migrating theatlantic.com from a sprawling PHP codebase to a Python application built on Django. Our first attempt at a load-balanced Python stack had serious flaws, as we quickly learned. Since then we have completely remade our stack from the bottom up; we have built tools that improve our ability to monitor for performance and service degradation; and we have developed a deployment process that incorporates automated testing and that allows us to push out updates without incurring any downtime. I will discuss the mistakes we made, the steps we took to identify performance problems and server resource issues, what our current stack looks like, and how we achieved the holy grail of zero-downtime deploys.

Summary

Frankie Dintino explains how The Atlantic reduced Django response times from roughly 80 seconds during deployments to about 150 milliseconds by combining monitoring, profiling, caching, query optimization, better hardware, and deployment changes. He argues that there is no single performance fix: teams should identify bottlenecks, take easy wins first, avoid neglecting frontend performance, and design deployments so workers warm up before receiving traffic. He describes using Nginx caching, A/B symlinked releases, uWSGI Emperor and Zerg modes, preloading, and cache-busting warm-up requests to prevent request queues during deploys. He also explains why ORM caching had little benefit for The Atlantic, whose database had low latency and was optimized for reads, while model instantiation itself was a significant cost.

Key takeaways

  • Measure bottlenecks with tools such as New Relic, Opbeat, Django Debug Toolbar, or Django Silk before optimizing.
  • Nginx proxy caching and carefully chosen cache headers can protect an application from crawlers, bad deployments, and temporary outages.
  • Replacing large numbers of hydrated Django model instances with targeted values queries significantly reduced request time.
  • A/B release directories and prewarmed uWSGI workers prevent deployment-time code loading from queuing and overwhelming requests.
  • ORM caching is context-dependent and may provide little benefit when database latency is low and model hydration costs as much as querying.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction and Performance Challenges Frankie Dintino describes The Atlantic’s Django performance problems and the improvements that brought response times down dramatically.
  2. 1:48 Performance Strategy and Monitoring The talk outlines a layered approach based on profiling, monitoring, quick wins, and attention to front-end performance.
  3. 3:21 Caching and Site Resilience This section covers CDN cache headers, proxy caching with Nginx, Django’s built-in caching tools, and protecting the site during outages.
  4. 5:46 Query Optimization The speaker discusses page and ORM caching, related-object loading, model hydration costs, database tuning, indexing, and hardware upgrades.
  5. 8:07 Request Queuing and Deployment Bottlenecks The talk explains how simultaneous worker startup caused requests to queue and overwhelm the application during deployments.
  6. 10:29 Zero-Downtime Deployments An A/B deployment strategy, preloading, worker readiness, cache warming, and uWSGI’s Zerg mode are presented as solutions.
  7. 12:48 Choosing uWSGI The speaker compares mod_wsgi and uWSGI, addressing documentation, configurability, community support, and real-world performance considerations.
  8. 14:21 uWSGI Monitoring and Scaling This section demonstrates Emperor mode, the uWSGI stats server, dynamic worker scaling, and the architecture used for multiple Atlantic properties.
  9. 16:29 Questions The audience asks about Django coverage, CDN traffic, bots, Gunicorn, and the practical benefits of ORM caching.

Transcript

3,494 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: Come on, no, yeah, yeah, yeah.

0:14

Speaker 2: Uh so today I'm gonna be talking about um basically the the performance issues that we dealt with at the Atlantic over the last year, uh how we overcame them, and some of the things that we learned on the way there. So I'd like to start off by thanking the organizers of DjangoCon for taking interest in my proposal, all of you for coming to listen to it. And I think one of the most valuable things about this conference, on par with the talks and the events, is the opportunity it presents to meet and talk to some really smart people. In fact, the work I did in the past year, what eventually became this talk, was thanks to a conversation I had with one of the founders of OPI. If you're not familiar with them, they're a service for monitoring performance of web apps similar to New Relic but with a special focus on Django.

1:02

Speaker 2: At the time, in September of last year, we were struggling with growing pains in our Django application. We had just completed a huge project, the porting of over six years of legacy code, a mix of PHP and Perl, to a Django-powered CMS in front end There were a few performance hiccups at launch, but we eventually overcame most of those, with one important exception. Our servers basically melted when we deployed. At the time we were using Mod Whiskey, so our process was basically to tee up the code using a fabric script and then touch the Whiskey file. Today, as you can see, and you're reading that correctly here, uh 80 second response times um and now down to 150 milliseconds. So um I'm gonna first cover the things uh there wasn't one single thing that we did

1:48

Speaker 2: that addressed our performance issues. It was a bunch of small things, a couple things that that were actually low-hanging fruit that like were great wins, but in general it was a bunch of small things that build up to fixing our performance problems. So the first thing is monitoring and profiling. If you don't understand what your bottlenecks are and you don't understand where the slowness is occurring or where um the app is freezing up um then you're just sort of stabbing in the dark Um and I also think it's important to tackle the easy stuff first um in quotes here. Uh query optimization and caching are sort of pretty easy ways to get quick wins with the application. And also um you know this talk is very focused on the server side of things, but it's important not to neglect front-end performance as well

2:36

Speaker 2: which is like a completely different ballgame, but you can have a super fast server and still your JavaScript takes 20 seconds to load and nobody visits your site. So for profiling, um we found uh the most use with hosted services, um New Relic or Opbeat are both comparable um and great services, and we identified a number of really important performance issues in our site with those. Also everybody I'm sure is familiar with Django debug toolbar. I recently came across Django Silk, which seems like a really promising application for sort of a different way of profiling your Django application. And if you're a masochist, um you can always try C profile, PyProf, 2Call tree, and Kcash

3:21

Speaker 2: Grind. And also there's some Whiskey Middleware which I've had mixed. experience with there was like one um uh bit of code that I was able to like kind of figure out where things were going wrong with that, but it was really difficult to set up and probably not worth the effort. Monitoring, there are downtime notifications, some of them hosted a lot of the same services that provide performance monitoring, New Relic upbeat, and then also alert trend chart beat. uh provide the service as well. And then you can have self-hosted services like Nagios and its many forks and also others like Zabbix. So with the easy stuff, caching, using a CDN, we at the time that we launched, we did have a CDN, but we weren't using our cache headers properly.

4:10

Speaker 2: And so kind of being really strategic about where we said what to vary the cache on and when and how long to cache the page, particularly for archival pieces. uh really helped us with some of the traffic we got from crawlers and bots, um which is the majority of the traffic that gets past our CDN. It's important, I think, to have a proxy cache, whether it's Nginx Squid or Varnish. We use Nginx and we've had uh A great success with it. One of the really great things about the the Nginx proxy cache is you can actually set it up to just cache for five seconds or two seconds, but to hold the information in the cache as long as they're until there's a 200 response. So if your application goes down, uh

4:57

Speaker 2: that two second cache key will hold up until your site comes back up again. We've had this really save us. Like, you know, somebody deployed some bad code or there was an issue with the database and the site, nobody really realized that the site was down. because we had this two second cache that basically lasted all night. Um also with caching um the stuff that comes built in with Django like the cache template tag, at cache property, cache static files storage to cache the information about where the static files are located on the file system, and the cache template loader, which does the same thing for the template finder. Page caching frameworks. We actually have one that I thought we had open sourced, and I realized today that we have it closed source for some reason.

5:46

Speaker 2: Um I talked to my coworker and we're going to open source it called Django CashCal in the spirit of sort of punny, stupid names for Django projects. It's kind of loosely based on Django Jimmy page, which is hasn't been updated in three years and still list as Alpha. So maybe it's an improvement on it. But there's also a number of really decent page caching frameworks. ORM caching like Django cache machine I feel like is a mixed bag. And I'll actually get to that in the next slide. So For query optimization, there's the obvious stuff like prefetch related and select related. Prefetch related objects, which is in 1. 10 public, though it's accessible in earlier versions of Django, that allows you to

6:34

Speaker 2: pass a callable and filter the objects that you pre -fetch on. so that you can prefetch conditionally items in the query set. For instance, in a generic foreign key, you can prefetch all of the instances of one model this field and not have to worry about it conflicting with other uh results in the query set that might be from different models that do not have that field. . I found this really surprising. When you're profiling Django, one of the things that we found to have like a huge impact but which doesn't really show up in any profiling is model instantiation and hydration because and this this is what ties back into like Django cache machine

7:21

Speaker 2: We had a query that basically we stored in the database. We had um all the different ad slots on our site And then the breakpoints that they were enabled and the sizes of the ad units. And uh altogether with the pre-fetrolating and the you know select related this amounted to I think like seven hundred instances being uh loaded into memory on every request which came out to about 80 milliseconds 100 milliseconds per request Um when we switched this to just dot values and you know manipulated a dick to basically simulate the prefetch related, like I said, 80 millisecond drop in every request. So that was like a huge gain, basically made our response times about 25 %. faster at that point. And then obviously, you know, but much more difficult

8:07

Speaker 2: database parameter tuning and judicious indexing. And you can always throw more hardware at the problem. That was um, you know, part of our uh speed gains there was cheating. Uh we upgraded from like four-year-old servers. to brand new servers that ran you know twice as fast. So you know that cut our response times in half. Now, uh you might notice from the earlier thing, um the only thing in this chart is request queuing. Um what is request queuing? So to explain that I'm going to rewind a little bit and just kind of talk about uh how we were set up. So we had um I think what's a fairly standard um way of load balancing a Django application.

8:53

Speaker 2: We had Nginx in front of it as a proxy cache, and then we had a couple of application servers, in this case they were running ModWISG behind it. The Nginx was listening on port 80, and then we had the different applications listening on different ports and defined upstreams in the Nginx config across the different app servers. We also had a stage version of our site, which we would use to test out code in production, on the production environment, and we just had like a different port for it. So, you know, our directory structure looks something like this. The discussion I had last year at DjangoCon with the founder of Outbeat, you know, so I was explaining him this problem we had. We would load everything up, we would try and preload as much as we could, then we would touch the WISGI file and then kind of cross our fingers.

9:44

Speaker 2: And if we were we happened to be slammed by a bot at that time, the site, all the code that it took to load, um On our local dev machines, just loading up one worker process, it would take maybe 10 seconds. But on the server loading 40 workers simultaneously across you know, five VMs which are all actually sitting on three bare metal machines, this could take thirty seconds, forty-five seconds, and during that that period of time, all these requests are queuing up. And uh basically the server, by the time the server's responding, it's still being overwhelmed by requests. And there were a few times where we had to do like rolling restarts of our server and actually bring them down and take them out of rotation in Nginx just that they could kind of cool off

10:29

Speaker 2: and and be able to respond again. So what he suggested was why don't you rather than having like a stage and a live that are separate and um you know doing the switch basically like at the deploy time load the code and touch the WISGI first and then um switch the simlinks. Uh basically have like A and B symlink to stage and live so that you have um The A and the B folders are always production ready. In any given case, one of them is either one version ahead or one version behind what's currently on production. But in every other respect, the you know The CDN path that they're using, the database, their

11:17

Speaker 2: production. You force as much code as possible to load before defining your def application in your WISGI file. This is less important for mod WISGI, but we switched to UWISGI. And uh one of the things that's nice when you're using UWISGI and you set up um again with the sort of silly names for uh Python and Django projects, um Zerg mode, where you have like um The way Zerg mode works, you have a Zerg server, which is really very bare bones. All it does is it just like listens for workers. And then the Zergs, when they come online, they communicate to the Zerg server that they are accepting requests. And it's only at that point that the Zerg server routes requests to them. So this gives them the freedom to sort of

12:03

Speaker 2: take as long as they need basically to get everything loaded up into memory. and um preload everything before it even needs to deal with the incoming requests. And the the Zerg server handles the queue. And then before swapping the upstreams and after they've been preloaded, we warm them up with a number of simultaneous concurrent HTTP requests. that basically make sure that all the workers get a request with a cache busted URL so that they're all primed. So why did we switch from mod whiskey to UWISGI? There's I think a lot of um misinformation about whiskey servers and um you know

12:48

Speaker 2: performance benchmarks and things like that. I think generally the comparisons and the configurations aren't very reliable And the thing being tested often doesn't apply to the real world. There's something like how many concurrent hello world responses can you serve up at any given moment? Mod WISGI in general, it's comparable performance, like I said. It runs on Apache HTTP server, which is, you know. Whatever. Mostly the work of a single developer. The documentation is terminally out of date and important configuration options are missing or only discoverable in the release notes, which is something that's openly admitted in the documentation. Kind of like the summary of the GitHub pages kind of demonstrates the the balance between

13:34

Speaker 2: where the community's gone with mod whiskey and e-whiskey. So UWISG broad and active community, very thorough documentation. In fact, it's kind of overwhelming, but also highly configurable. And So when I was rehearsing this talk, there were a couple slides that I don't think I'm going to have time to get through. But I do have a GitHub repository which uh has some demo code that kind of demonstrates how we're using uh UWISGI Emperor mode and Zerg mode, and the UWISGI stat server to kind of monitor our WISGI applications, preload everything using a fabric script. So, you know, I'll just sort of reference everyone, uh

14:21

Speaker 2: refer everyone to that and sort of as a teaser kind of show uh what are UWISGI monitoring page looks like. So it looks something like this. We use Emperor Mode. Emperor mode basically is a way where you can uh dynamically uh set up multiple configurations for a single code base. So we have like one really large code base that hosts all the Atlantic properties, City Lab, The Wire , the Atlantic. com. We have a profile site which is for people to manage their print subscriptions. We have a sponsored section of the site, and we have an A and a B version of all those things. And then also we have a CMS that's for our editors and a CMS called Waldo that's for outside video contributors.

15:06

Speaker 2: So uh we go to this page, we can see uh in real time as requests come in. Um There we go. It sort of changes. You can sort of hover over the uh uh little dots to see information that's pulled from the USC stats server If the site starts to get overwhelmed, there's a little queue monitor down at the bottom. Generally, if the queue fills up a little bit, that's normal, uh that's sort of expected. And we have uh one of the things you can do with the Zerg mode is um dynamically spin up instances based on how busy the other worker processes it are. So you know when these requests come in,

15:51

Speaker 2: spawns new processes and then it's able to handle uh the in you know the increased uh volume that's coming in So up here, these are the URLs to the GitHub repository. I actually haven't pushed it yet, but by the time this video is online, it'll be there. I think it'll probably be sometime later this week. At Frankie Dentino, I don't really use Twitter much, but it's an easy way to get a hold of me, or you can email me, Frankie at theatlantic. com. And with the time I have left, I'll open it up to any questions.

16:29

Speaker 3: Great talk. Thank you. So first question, or just the question is, uh is basically all uh end user traffic that hits the Atlantic. com served Uh like through C DN but then that C DN is basically talking to Django across the board. So is the entire site served off of Django, I guess is the high level question

16:46

Speaker 2: Yeah, well uh mostly. I mean there are a couple of really really random legacy pages that look really old and they are really old um but 99% of the traffic goes through Django.

16:57

Speaker 3: And like homepage and like volume.

17:00

Speaker 2: Yeah. And um as far as the CDN is concerned, so a small portion of the requests uh will get through the CDN. You know, I we generally have about a three-minute page cache for every page, and then we have longer ones for some, you know, older pages that haven't been updated in a while. But most of the requests that get past that are actually bots and crawlers. We have an archive, I mean we've been printing since 1857. And we basically have everything online that we have the rights to. You would think that a print magazine would have digital publishing rights to all their articles, but you would be wrong. That is not the case. But everything that we do have the rights to and that we've digitized is free online. So you know, we don't generally have an issue when people crawl our site if they're doing it in good faith, but occasionally there's like an AWS instance

17:48

Speaker 2: that'll make like 200 concurrent requests every second and uh sort of makes our servers melt and so we have to It's kind of scramble to block it. And so a lot of this was like kind of dealing with that and getting us in a place where we could handle those sorts of situations without breaking a sweat.

18:04

Speaker 3: Great, thank you.

18:08

Speaker 4: Hi. Uh so you talked about using mod whiskey and switching over to eWISKY. There's a third option that a lot of people use. I was just wondering if you thought about G Unicorn, didn't go with it, or just didn't have time and you had seen how big UWISGI was. Do you have an opinion on that?

18:25

Speaker 2: I don't actually. I know people have had a lot of success with it. I don't really have a strong opinion one way or the other. Just kind of you whiskey worked for us, but I think uh I think it's it's on from what I understand G unicorn G Unicorn is on par with U Whiskey in terms of features and performance.

18:43

Speaker 4: So you didn't you didn't like look at G Unicorn and decide like, oh, that's not going to work for us.

18:47

Speaker 2: No.

18:47

Speaker 4: Okay. Thanks.

18:49

Speaker 2: Sure.

18:51

Speaker 5: Yeah, you mentioned earlier that uh you thought um orm caching was kind of a mixed bag. Can you go into more detail why you think so?

18:59

Speaker 2: Sure, yeah. So uh when we uh we ran Django Cache Machine for a really long time on our site and then uh actually kind of accidentally we turned it off. Um and And and this wasn't the first time we've accidentally turned off caching. But in this case, there was really no performance difference. And we kind of speculated about why that might be. Um and uh the conclusion we came to is that A lot of the like so our situation is maybe unique because we're not in the cloud, we're in a data center, it's all you know in REST in Virginia, all like connected via fiber channels, so there's not much latency between the servers and the database. And the database is pretty well tuned for reads because we don't have a lot of writes because we're, you know, the editors publish maybe 20 stories a day versus people reading millions of stories a day.

19:49

Speaker 2: So the database is generally not a bottleneck for us. and uh the model instantiation and hydration and basically the unpickling and creating of the model instances was taking almost as much time as if we were doing it raw ORM queries uh to the database. So it w like really didn't make a big difference. But I think that in cases where you're like on a cloud where there's latency between your application servers and your database servers, or uh when you have situations where you're like writes and reads are more balanced, then I think there might be a uh like I said, a mixed bag, then it might uh be useful.

Questions this talk answers

How did The Atlantic reduce its Django response times from 80 seconds to about 150 milliseconds?

There was no single fix: they combined monitoring and profiling, query optimization, caching, better deployment and process management, and newer hardware. Small improvements, including replacing expensive model instantiation with `.values()`, added up to a major reduction in response time.

Discussed at 1:48

How can an Nginx cache keep a Django site available when the application goes down?

The Atlantic used a short cache period, such as two seconds, but configured Nginx to keep serving the cached response until it received a successful 200 response from the application again. This allowed the site to remain available through bad deployments or database problems.

Discussed at 4:10

How can Django model hydration slow down requests, and how can you optimize it?

Loading and hydrating hundreds of model instances can consume significant time even when ordinary profiling does not make the cost obvious. The Atlantic replaced that work with `.values()` and manually shaped dictionaries, cutting roughly 80 milliseconds from requests and improving response times by about 25 percent.

Discussed at 7:21

How can you deploy Django code without causing requests to queue up?

They used A/B production-ready directories, loaded and preloaded the new code before routing traffic to it, and switched symlinks only after the workers were ready. With uWSGI’s Emperor and Zerg modes, workers could finish loading before accepting requests, and concurrent warm-up requests primed them before the upstream switch.

Discussed at 10:29

Why did The Atlantic switch from mod_wsgi to uWSGI for Django?

The main reason was uWSGI’s active community, thorough documentation, and extensive configurability, including the Emperor and Zerg modes used for deployment and worker management. The speaker said mod_wsgi’s performance was generally comparable, but its documentation and configuration support were weaker.

Discussed at 12:48

Why is Django ORM caching sometimes a mixed bag?

ORM caching may provide little benefit when the database is nearby, read-optimized, and not a bottleneck, because model creation and unpickling can cost nearly as much as querying the database. It is more likely to help when application and database servers have network latency or when reads and writes are more evenly balanced.

Discussed at 18:59

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Frankie Dintino

More videos from DjangoCon US