Creating an Inclusive Django Community with Kenya Phelps
Published July 15, 2026
This video features Ganesh Swami at DjangoCon US 2017 in Spokane, Washington, USA.
DjangoCon US 2017 - Butter smooth, interactive applications with Django and Websockets by Ganesh Swami
Web applications have changed significantly over the years – from simple static pages, to sprinkling interactiveness with JQuery/AJAX, to full dynamic single page apps. Through each evolution, we’re adding more complexity, more data and more asynchronous behavior to our applications.
In this new world, where does the synchronous nature of Django’s request-response cycle fit in?
My talk will focus on the topics around asynchronous Django applications. I’ll be sharing some lessons we learnt while building and scaling an interactive web application within the confines of Django and django-channels.
This topic is interesting because there’s been a lot of interest with meteor-like frameworks that have synchronized state between the frontend and backend. My intention is to show the audience that you can accomplish the same end-result with Django, without the need to learn and deploy a brand new framework.
An outline I have in mind:
What does asynchrony mean, and why you need it.
Traditional methods of achieving asynchrony (delayed jobs using worker queues like celery, long-polling for messaging, etc.)
Why django-channels changes the game.
How to architect your state.
What are the available options for deployment.
Gotchas, and what to do when things go wrong.
Just a basic knowledge of Django is required, as the topics are transferable to other frameworks. We did not have to monkey-patch any of the drivers to achieve asynchrony, so what you’ll learn at my talk will apply cleanly to a stock Django.
This talk was presented at: https://2017.djangocon.us/talks/butter-smooth-interactive-applications-with-django-and-websockets/
LINKS:
Follow Ganesh Swami 👇
On Twitter: https://twitter.com/gane5h
Official homepage: http://www.silota.com
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Ganesh Swami explains how his Django-based SQL analytics product moved from ordinary Ajax request/response handling to a hybrid architecture using Django Channels, WebSockets, Celery workers, and Redis. Long-running user queries, sharp dashboard traffic spikes, and deployments that killed active queries made synchronous requests impractical. The new design lets Django schedule work in the background and notify clients when results are ready, while keeping the existing Python and Django stack. Most of the presentation covers production concerns rather than basic WebSocket usage: choosing a compatible AWS load balancer, recovering missed messages after disconnections with ordered messages and Redis sorted sets, routing events through a front-end event bus, and testing by replaying recorded message streams. Swami argues that interactive applications require explicit state management, reliable message recovery, and careful deployment and debugging practices; WebSockets solve only part of the problem.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Thanks Mark. Can you guys hear me good? Awesome. Excited to be here. Django Con, uh my first Django Con. Um besides the people who are always Python community is A plus, the food was Exceptional. I think a lot of people a lot of you guys are still in the food coma after lunch, but let's take it easy. So the title of my talk is uh ButterSmooth Interactive Applications with Django and WebSockets So I'm going to be giving a little preview of uh interactiveness on the internet with regards to web applications. and then build up to its Django channels and WebSockets and uh the kinds of problems that we faced. I'm not going to go too much detail into
Speaker 1: how web uh Django's channels work because it's a work in progress and there I believe there was a full tutorial on Django channels and uh a lot of the other talks uh talk about Django at channels specifically. So I'm going to be talking about the tooling and some of the deployment issues that we've had A more uh formal introduction about myself. I've been uh working with uh Python data for over ten years. Uh started my career writing assembly programs. programming and just work my way up to up the stack. So I'm like a full stack dev but stops at the at the server level. So I've written programs for GPUs, uh, uh a lot of that kind of stuff.
Speaker 1: Uh but I I really love Python, I love Django, I love how clean and uh none of the magic, none of the uh black box stuff you can just open up the code and see what's going on. I've been doing data analysis in SQL for the past couple of years. My claim to fame is I wrote the first wiki engine for Emacs. and uh wiki blog engine so uh that was pretty cool it was accepted into mainline uh mainline emacs and this was a long time ago. I currently run Solota, which is uh is a is a data analysis tool for professional data analysts, um SQL recipes which Our product is something like this.
Speaker 1: This is just to give you a context on the challenges we faced. You type in SQL there, it's a text box. And you hit run and it creates these charts. And then you can schedule these charts, share these charts, and do whatever you want. So that's essentially our product. Who uses our product? Professional data analysts. So if you're uh it's not SQL like the ORM style SQL, it's for building histograms, correlation. uh doing forecasting, regressions, all the kinds of data analysis stuff that you would do, but using SQL. So if you Google for Siloda SQL recipes, there's a full list of I think two dozen recipes You can copy paste, everything is straight up SQL.
Speaker 1: So this is a timeline of uh asynchronous behavior on the web. Uh how many of you guys are familiar with iframes? That's awesome. So about 20 years ago, iframes was how you'd be uh you get interactive uh web applications. If you guys are familiar with MapQuest and some of the I guess the pre-Web 1. 0 kind of applications, you click a button and then it dynamically generates iframe URLs and then it refreshes. I believe in even Hotmail, the first version of Hotmail worked that way. So you click send and then it refreshes its iframe. So that was like interactive web apps. And then around 1999, Microsoft introduced XML HC.
Speaker 1: CDP request, which is the precursor to Ajax. It's a way of creating asynchronous request responses. Around 2004, Gmail launched, and Gmail was the first, I would say Kind of a single page app, so there's no full page refresh when you hit send to send an email. Around 2005, Ajax, the term Ajax was coined. Uh 2005 was when Django was released publicly. And then 2012, Meteor JS. How many of you are guys are familiar with Meteor JS or Herdiv? Awesome. So Meteor JS is like a was mind-blowing when I first saw well how it worked.
Speaker 1: So it was like an interactive fluid web framework where there was a subset of your database being held on the front end. So you query your front end and it's it's pretty fluid, it's pretty uh amazing how it all works. I I still don't know if it's a good idea, but uh it's pretty sweet. 2014 React was public, started gaining adoption. I believe that's when the single-page app acronym SPA was uh started getting uh into the main Mainstream 2016. I believe there's a small framework called Rails. They now include WebSocket support into their mainline. They call it AC Action Cable.
Speaker 1: So that's roughly the outline of interactiveness on the web. So this is a a very common way to visualize how technologies grow and get adopted across a wide spectrum. So I would say pre-2004, a lot of innovative webs, a lot of risky applications were being built. Probably only 10 or 15% of the browsers out there supported the technologies. But those were that's the innovative uh innovation and interactiveness being introduced. 2004 to 2014, early majority, big teams, a lot of resources, very expensive to build these interactive web apps But it was being done.
Speaker 1: Case in point, Gmail, Google Maps, a lot of the first version of the web streaming services, and so on. 2014 and up to now I believe we're uh inching towards the promised land. It's uh a lot of the best practices have been established. Uh there 's still some challenges on the fringes, but more or less we understand how to build a single page app. What are the consistency What are the difficulties? There are at least three or four mature front-end frameworks that work out of the box. Vue. js is a popular one. React. js is another popular one. A lot of folks use Ember JS, so that's pretty awesome. So I think we're getting there.
Speaker 1: Between 2014 around then, there have been a lot of the commercial applications of these fluid uh you know push notifications and that kind of stuff. So socket I. O. pusher, Firebase. Meteor JS, these are all a hybrid of uh open source and uh and uh backend as a service kind of system to have uh interactive web apps. So the main question is why do you need this? Like what's uh what's the need, right? Because Django has been around for ten plus years And uh it works pretty sweet. Uh why would you want to change something that's
Speaker 1: that's working perfectly fine? So instead of answering this broad question about why, I'm going to be uh scoping it into our web app and why we saw the need for WebSockets and interactiveness. and why uh just regular web apps didn't cut it. So as I mentioned, our product is a is a SQL editor, essentially. You have a text box, you hit run, and then it creates these charts. And then you go through this workflow over and over and over again. So that's basically what our product does. So if you were to draw a schematic, On how this works, you have the app, which is the web app, running the browser.
Speaker 1: There's an Ajax request that reaches the backend And then the backend, which is Django, executes the SQL query and then returns a HTTP response. So that's basically the lifecycle of your one one single iteration If there is a mistake in your SQL, your execute SQL throws an exception and then it's sent back as a HTTP response with a right error code. And your friend handles that. So this this worked for you know version 1. 0 when we had no users. Uh everything worked great. So it's it's always pretty fun to see you hit run and then it It works. Okay, so awesome. So let's go to production. The minute we hit production and once we got uh a couple of users, we started seeing these problems that were creeping up.
Speaker 1: So I'm just going to enumerate some of the problems we've had with this regular uh request response cycle. The first problem is unbounded runtime for the request. So here essentially what's happening is we let the user type in arbitrary SQL So we do not know beforehand how long this SQL is going to take. It could take you know three minutes, it could take 30 minutes, it could take half a day, nobody knows. Because it's up to the user what what kind of model or what kind of analytics they're running. So the problem was some of our users were typing their SQL, hitting run, and then closing their laptop because they knew it's going to take half a day to run the query.
Speaker 1: And from our side, you know, having that connection, that request response, coming back and restoring the state was being uh was a big challenge. So that's the first difficulty. Uh the second difficulty was uh what is this what we call 8 a. m. traffic spikes So in our product, you can organize the SQL queries into a dashboard. So you can have half a dozen to a dozen different charts and create a dashboard. And at 8 a. m. our users load up these dashboards and you have 15 queries hitting our backend and they're all trying to run. And so the the ratio between the peak load to the low point was about 7x. So you had to provision 7x
Speaker 1: the number of servers and resources to handle that load. Um so when I casually talk to my non-technical friends, Um say, yeah, I was at the office at eight AM It's like, Oh good for you. You wake up at six, hit the gym, go for a run, make your breakfast and then you were at work at eight A. M. No, actually no. I uh I get paged at seven fifty five that the servers are down and then I rush to the office to fix. them but I don't tell them I just say yeah I hit the gym and you know do all that stuff. Uh the third challenge was about deployment because the Django servers were actually ex Executing these queries, you couldn't really swap out new versions of the code. So if you have, let's say, five servers that are executing these uh SQL queries and they could take anywhere from 30 minutes to uh let's say uh
Speaker 1: three hours. Deployment was a problem because deployment would kill that executing query and the user would never see the results. So, you know, we believe in active deployments, quick deployments. The minute the unit tests change uh or they succeed, they just push into production. Uh but we were not able to do that because our users queries were failing or uh were being terminated. So that was a big challenge So that is a summary of the challenges we've had with the synchronous request response cycle. So we had to look for an alternate solution. So looking around, you know, um I think I met Andrew last year or the year before, and he was uh just starting out, maybe this is like version point zero zero one of Django channels and I was pretty excited about it.
Speaker 1: It's like okay Jango Django channels seems like the holy grail right what can go wrong so let's uh let's play with it. And so High level how Django channel works is that you still have your synchronous request response cycle, which is on the on the top half. And then you have a persistent bi-directional connection based on WebSockets in the lower half. And so you send a message to your back-end through the Ajax request And then the backend sends in the response through the HP response, just as usual. But at the same time, you can send messages back and forth through the WebSocket. So your server can initiate a message without having a matching HTTP
Speaker 1: request or Ajax request. So that's conceptually How WebSockets or Django channels work. They don't replace anything, it's just an add-on to your existing workflow. So the the method we used was you continue to so when our users hit run it still makes a HTTP request to your back -in. And then your backend, which is Django, offloads it to this worker queue, executes the SQL, and once it's done, using WebSockets notifies the front-end that the results are ready. So we're able to push it out of band.
Speaker 1: So the uh just as a thought experiment, the alternate old school approach would would have been uh using a thread on the on the front end to pull your back end for results and uh if there were results you get a response uh but then that's undue load, right? You can't set the polling interval based on the query because you don't know. So you're going to hit the backend, let's say, every 30 seconds. And that's going to break if the user queries only take two seconds to run. So the long polling approach had its own issues. So uh pretty amazing. Just went to the documentation, copy paste the examples into our code, uh
Speaker 1: two days, and we had it all working. At least it was working Everything was working on my laptop. Okay, uh amazing. Let's go out for drinks because this project is done. Um, you know, where can I sign up for my bonus check? Uh well, no, not really. It took us about six weeks to make it production ready. Uh so going from our laptop to our cluster of servers uh took a lot of time. So you can copy paste examples from the Django channels, uh but the rest of my talk is gonna go into The challenges we faced taking it to production, which I g I guess is the meat of my presentation. So it it works on our lap on our on my laptop, you know, an MVP. It's a proof of concept.
Speaker 1: Uh the challenges seem to, you know, it looked promising. So this is when you sit down and write a business case. for how to re-architect your app because this is pretty expensive. You're writing the core or rewriting the core of how you're doing work. So we came up with this uh a list of bullet points, uh, and I'll just walk through them. So the the big overarching aim of of this project was to make the back end the driver of work. The front end, just because the user says run this, uh we're not going to obey that command. We're going to schedule it and work on that query when we see when we see is as a good time. So we are going to be the back end is going to be the driver rather than the client.
Speaker 1: Sane handle handling of state. You know, the uh the state could be a query is in progress, the query has failed, the query has succeeded, the query is ready to be run, uh, and then there are disconnections in the middle because people People could go to trains, people could shut down their laptop. So you just have to map out all the possible states and try to understand what's going on. Just a clear picture. Server side events, there are lots of approaches for doing this, but the server should be able to communicate new data to the front end without the front end asking for it. I think that's a very powerful concept. No proprietary solutions, not that I'm against propriety solutions, but I feel that when you're working with brand new technology, it's a lot harder to debug propriety
Speaker 1: technology So solutions. So uh one of the difficulties I've had with um socket. io is you had to run socket I. O. on the back end, uh which is Node. js, and uh it was pretty challenging to uh debug what was going on uh because i come from a python backend and debugging an external stack was was a challenge Uh keep investments in Python plus Django stack. Um I guess what a lot of other people would have done is they would have just thrown in Meteor. js or a Node. js stack and then offload this asynchronous component to Node. do to Node. js and let Node. js deal with this because there are lots of frameworks on Node. js that help you do this. We do not want to do that
Speaker 1: because every time you have an architectural decision and you add a new layer to the stack, I think your uh your product becomes kind of a Frankenstein product and it just It's just too much to keep in your head because now if you're trying to hire someone, they you're asking them for like React front-end experience, you're asking them for Django experience, you're asking for a Node. js experience and then the deployment options. It's just too much. So we wanted to keep everything standard out of the box, Python plus Django. Off-the-shelf AWS components, our stack is completely on Amazon Web Services, so we do not want to uh build something custom, uh which means no recompiling engineering Next to add uh WebSocket support
Speaker 1: uh because we wanted to use Amazon's load balancer, for example, because it has a lot of benefits. It gives you health checks, it g has good logging uh a good logging framework. It does DDoS attacks, a lot of lot of good features, and we wanted to tie into that ecosystem. So it didn't make sense, a lot of sense for us to bring in our own infrastructure components. So roughly this is uh this is our architecture diagram. So we have our clients on the bottom, which could be uh I just found this in some kind of stencil program. There are it's a phone, it's a laptop, and it's a desktop. So three kinds of clients. uh and they connect to a load balancer
Speaker 1: which has web socket support and these uh the load balancer then starts to load balance to a back end which could be the WebSockets plus the regular G unicon. So it's a it's a hybrid it with the synchronous and the asynchronous part. And then Redis is the persistence layer where the channels are being uh shared. So if you go through the Django channels documentation, this is one uh approach that's recommended using Redis as a backing server. And of course you still have your standard uh Django components, you have a database, you have your ORM and all that stuff. But roughly this is the architecture. So now the challenges. Uh why did it take us six weeks?
Speaker 1: You know, uh that's that's a little odd. You know, did I go on vacation for four weeks? four weeks during the project? No, not really. Because uh we were iterating through the solutions and we found that something was wrong and then we had to go back to the drawing board and then we were trying to swap out the stuff. I think there was some conflict with our initial motivations and our business plan where we wanted to stick to just the Python stack. So uh it was a lot lot more effort to get it it to work. But I think in uh now I'm really happy with uh how it's all working out and going forward it's going to be a breeze to maintain because there's no custom code anywhere. So the first challenge was cloud deployment. I'm gonna go step by step into the challenges we faced.
Speaker 1: So uh if you remember the architecture diagram, we need a load balancer in front of Geunicon and the asynchronous website. server. The problem with a lot of these standard off-the-shelf web uh load balancers is that they're HTTP load balancers, uh but what WebSockets needs is a TCP load balancer. So all the stuff if you're familiar with Django's cookie base authentication, cookies are a HTTP layer concern. So if you have a TCP load balancer, you can't really attach that cookies, like it can't read the cookies. So if you're trying to have sticky sessions based on cookies, it's not really going to work. And I was surprised to see this, but a lot of the load balances out there on the cloud were all HTTP
Speaker 1: load balancers So Google Cloud, Azure , Amazon. But fortunately, Amazon launched a brand new load balancer last summer called the Application Load Balancer, which is a TCP-level load balancer. But uh The problem with a lot of Amazon's products is that the documentation is just straight out wrong or uh they say yeah we don't recommend you doing this, but that is exactly what you need to do. So I was kind of fortunate that I was taking the train from Vancouver, BC, which is home for me, down to Portland for PyCon, and uh one of the uh elastic beanstalk engineers got on the train, so I just fed him like tons of beer and ask him like why the hell is this not working? What's wrong with this? Why is your documentation wrong?
Speaker 1: Why do you have three versions of the same stuff which are all like in different layers of compatibility. So the lesson here is make sure it's working with the application load balancer. The deployments on your laptop are very different from the deployments on the cloud. uh and try to try to figure out how application load balancer works with your rest of your stack The second challenge we had was uh to do with missed messages on disconnect. So here's an example. Let's say Django periodically sends a server send message to your front end, right? So it's moving upwards. And so the green bars is when your client is connected to the back end. And then for some reason it is disconnected.
Speaker 1: Either the user goes through a tunnel or there's some kind of disconnection notice and the server sends a message. But this message has nowhere to go because the client is not connected. So that's the stop sign or the no entry sign. And then the client reconnects and it's missed those messages. And then it happens again. So if you're going to be missing a lot of the messages that your backend sends you, and the back end has no no kind of system to uh to acknowledge that these messages have been received, you're gonna be left in a very bad shape. Your s your client state is just going to be in an invalid state. It's it's pretty challenging. And uh I was surprised to find that on the cloud it's actually quite frequent to disconnect. Uh persistent connect.
Speaker 1: There's just not a guarantee. Uh there is an NPM package called Reconnecting WebSockets. Essentially does a for loop, and if it's a disconnect it uh tries to reconnect. So that was easy. But what we had to do is actually fix the missed messages on the back end So our solution to that was to serialize all of the messages being sent from the backend, timestamp it, and implement a cursor on the client side. So if the backend sends message one, message two, message three, Message three, message four, the client keeps track of one, two, and then it receives four. Okay, you know what, there's a problem. So go back to two and give me all the messages since two. So that's roughly how it's essentially a table that we store
Speaker 1: on the back end and then the client implements this logic where if there are if there's a gap in the sequence it can go back and rewind that cursor. So what we uh did here as an implementation level detail is not to store it in our database, but uh Redis has this data structure called sorted sets. And so you can actually pop and push into this Reddit structure and it's actually really fast. So there's no uh there's no delay in uh in trying to store this this uh ring buffer. We also wrote a middleware because we wanted to make sure that every message going from the backend was going through was being timestamped and
Speaker 1: sequentialized. So it was important to go through that single stream. Our uh third problem was uh front end architecture. Our front end is uh is a single page React app. And these messages from the back end, you don't really know when they're going to come. They could be out of order. It could be different components. Because it's a single bus that all these messages come in. And so we were uh actually pretty confused in terms of how do you update your front-end state. Uh and so what we did there This is an example. Let's say that's your web app and you have a notification system that gives you a drop-down with all the new messages.
Speaker 1: And then you have a feed which tells you all the new elements that are inserted into your feed system. And our WebSocket handler gets all these messages, and then you use standard jQuery's uh on messaging. So you essentially implement your own messaging system on the front end to dispatch these messages. So if you guys are familiar with uh how iOS apps are built, there's a single bus and you get these messages and then they're dispatched to the different different views and these views uh listen when they load it and then they destroy the handlers when they unload. So it's the same kind of concept. That way your notifications system and your feet are not tying directly into WebSockets, there's an abstraction layer there. And for whatever reason, if WebSockets is not supported,
Speaker 1: you can always fall back. This is just an architectural idea on how to organize your front end. Problem number four, debugging. Debugging, testing, all of this was a was a big challenge, uh, because you cannot use curl to mark requests. A lot of the Django test clients and all that stuff doesn't work. So what we did is uh you know that example with the middleware that serialized the messages? We just stored that as uh as our fixture. So we were able to bring up a front-end client and then feed it messages from the back end based on this fixture. And then it's pretty important to keep that stream of messages on your back
Speaker 1: end and the stream of messages received on the front end and push them both into a centralized data store. So you can reconcile the messages that have been sent from the back end and the messages that have been received from the front end and see if there are any gaps. So without this unified view across the client and the server, uh it's kind of difficult to figure out what's going on Here's a another tip. Chrome has this pretty awesome WS tab. I spent two weeks debugging Chrome WebSockets without knowing Chrome had this built-in. So you can just click on the frames and this gives you a sequential list of all the messages received through the WebSockets. It's either outgoing or incoming, no problem. It just shows it to you.
Speaker 1: And uh we just use JSON as a transport So you can uh you can inspect it, you can prefy it, whatever you want. So guys, that concludes the my talk. And uh I am open to questions.
Speaker 2: Thanks for the talk. Um, in one of your earlier diagrams, it looked like you were basically uh accepting the initial HTTP request via Ajax, running the SQL query like on a different Django box on a worker somewhere. Um, how are you able i is that true, first of all? Um were you doing it like in celery in the background or Yeah, that one. There we go. Right, are those two separate Django web servers?
Speaker 1: That's exactly right. So we were offloading the query to a salary worker.
Speaker 2: So how are you able to basically keep the web socket alive but when you move it from the initial machine that got the request to the celery web server?
Speaker 1: So the uh Django channels documentation recommends Redis as a persistent layer. So from any web server You can talk to a client no matter what actual web server it's connected to. So it gives you a unified view of all of the clients connected
Speaker 3: You mentioned that you use the an event bus on the front end with React. Uh did you did you use something uh like Redux or what what was the did you hand roll one for for doing that?
Speaker 1: We uh just use something uh built on jQuery, jQuery event emitters.
Speaker 3: Okay. Um
Speaker 1: yeah.
Speaker 4: I have a question. Yep. Um you mentioned um sort of the the hiccup in the original deployment, uh when you were just using Ajax request response. And I wonder how how you handle sort of that in this architecture where you still need to, you know, replace the code in that worker that's running the query. Ha how do you handle the deployment deployment of of the worker um that that maybe is still doing work.
Speaker 1: Very good question. So um our deployment cycles are different for the Django app and the salary workers. The workers uh it's okay to run with let's say v1 of the code and then you can actually do uh delayed deployments all new workers have version two and version one is still running, no problem. It's not, you're not killing the existing worker. So one worker, one instance of the worker is one query. So it's a lot easier.
Speaker 5: How do you get the static data in there? Do you just use do you just do a regular ETL to a Postgres database and then have them as unmanaged? Uh, you know, like the meta unmanaged Django models so that there's a connection to the front end. Um Uh did that question make sense? So you must be running SQL on some static type of data, right? Like if you're doing data analysis, you know, could be gigabytes of data from flat files. you load it into Postgres, then do you somehow introspect and generate uh unmanaged models? And and the reason I say unmanaged is because you don't want to Uh y
Speaker 5: you don't want to spoil any of the static data and accidentally have um you know what I'm saying. But unmanaged is confusing, but it it's it's the meta tag for the model to not have it migrate beyond the
Speaker 1: Very good question. Okay. So the question was, we execute the SQL. against some unknown schema. And his question was, do you import that schema into your Django models or does it pollute your clean, sanitized, uh curated schema or not, right? Right? Because it could be unknown. And uh the answer to that question is we don't manage any of our customers' data. They don't even touch our Django, our Django knows uh knows nothing about it. And so we literally send that raw text in that execute SQL box through a SSH tunnel to the customer's database which they host. So we don't we don't import any models, we don't do it's nothing to do with Django. Our customers don't even know uh we're using Django in the back end
Speaker 1: So there's a clear separation between our app and the customer's data model. Does that make sense?
Speaker 6: How did you solve the authentication issue? Maybe you uh answered it and I just just didn't hear it, but you said that we weren't sending the the cookies over and then you kind of like just maybe jumped ahead and
Speaker 1: it's uh there's a full section in the the Django channels documentation on authentication, but essentially what it does is it passes off the HTTP session to the Django session uh channels session and that handoff is done during And so the subsequent messages don't have any kind of uh cookie header or anything. So but the Django channels documentation goes into a lot of detail there. So uh you just have to add some decorators to your channel Views essentially.
Speaker 6: You mentioned that one of the reasons you wanted to adopt the solution was um Uh the n the 8 a. m.
Speaker 1: Yeah
Speaker 6: uh server load other than the removal of the long polls and are you ready yet type messages, did this solve that problem? In any other way?
Speaker 1: So we uh we definitely partially solved the problem in the sense that we still have that big spike of Ajax requests. But the difference now is that That full life cycle there is less than 50 milliseconds. It's not a long-standing connection. So you can connect 7,000, you can open up 7,000 connections that are 50 milliseconds each rather than having 7,000 connections. which are three minutes long. That's the difference.
Speaker 7: You mentioned you're using TCP load balancing and there so in in the TCP I mean you're so essentially are are you able to get the same uh level of affinity that you would i in sort of a a higher level scheme for load balancing such as HTTP and cookie and session based affinity.
Speaker 1: That's been a challenge for us. Sticky session is an unsolved problem in this scenario. Another challenge that we have is the IP addresses of the clients key. changing and so you can't really reconnect to the same uh uh web socket connection. So I think that's the reason why there are a lot of disconnects and reconnects. Uh I haven't gotten to the bottom of that yet, but that is definitely an unsolved problem.
Speaker 8: How do you manage uh lifetimes of connections? Because I can imagine that the reconnection um example that you showed could look very similar to a user saying, oh, I'm done with your website. Um I don't I don't actually want any updates.
Speaker 1: Which which portion, which lifecycle, from the web app to the backend or the backend to the Customers database.
Speaker 8: So the c the web socket connections. If uh if they go away, they can be intentionally going away or it could be a connection issue.
Speaker 1: So when they go away it unloads the app, so we disconnect the web socket. Right now there's no way for you to dis go offline essentially, right? There's no way to go offline. pretty aggressive in reconnecting and trying to reconnect. We have an exponential back-off there just in case the back end is actually down. Then you You don't want to like hammer your backend with every trying to reconnect every 50 milliseconds or 500 milliseconds. So you back off 500 one second, two seconds four seconds just like Gmail the Gmail you see it's reconnecting in three seconds and then if it can reconnect reconnecting in seven seconds and then reconnecting twelve seconds reconnecting twelve tomorrow and so on, right?
Speaker 1: So it's the same kind of a deal.
Speaker 4: Thank you again for sharing your expertise.
Speaker 1: Awesome.
The normal Django request/response approach struggled with queries that could run for minutes or hours, large dashboard traffic spikes, and deployments that killed queries already in progress. Django Channels lets the request enqueue work and the server notify the browser asynchronously when the results are ready.
Discussed at 9:39They do not replace HTTP: the browser still sends the initial request through the normal Django application, while a persistent, bidirectional WebSocket connection provides out-of-band server-to-client messages. In this application, Django queues the query and uses the WebSocket to announce completion.
Discussed at 12:50The application uses an AWS load balancer in front of synchronous Django and the WebSocket server, with Redis as the shared channel layer and the usual database and Django components behind them. The speaker emphasizes using an Application Load Balancer and verifying how it works with the rest of the stack, since ordinary HTTP load balancers are not sufficient for WebSockets.
Discussed at 19:45The backend serializes and timestamps messages, while the client maintains a cursor for the last message it received. If the client detects a gap after reconnecting, it asks for all messages after that cursor; the implementation stores the message sequence in Redis sorted sets.
Discussed at 24:28The WebSocket handler acts as a single message bus and dispatches events through an abstraction layer to the relevant UI components, such as notifications or a feed. Components register and remove their handlers as they load and unload, which also leaves room for a fallback when WebSockets are unavailable.
Discussed at 26:03The team records serialized backend messages as fixtures, feeds them to a frontend client, and compares sent and received message streams in a centralized store to find gaps. Chrome’s WebSocket frames panel is also useful for inspecting incoming and outgoing messages.
Discussed at 27:37The query is offloaded to a Celery worker, and Django Channels uses Redis as a shared persistence layer, so any web server can communicate with the client regardless of which server accepted the original request.
Discussed at 30:23Django application deployments and Celery worker deployments use separate cycles. Existing workers can continue running version one while new workers run version two, because each worker instance handles one query rather than being killed during deployment.
Discussed at 31:36It transfers the HTTP session into the Django Channels session during the connection handshake. Subsequent channel messages therefore do not need to carry the browser’s cookie header, using the authentication decorators described in the Channels documentation.
Discussed at 34:15It only partially solves it: the burst of Ajax requests remains, but each request lasts under about 50 milliseconds instead of holding thousands of connections open for minutes. This greatly reduces the duration and resource cost of the spike.
Discussed at 35:07Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026