Pessimism, optimism, realism and Django database concurrency
Published October 8, 2023
This video features Aivars Kalvans at DjangoCon Europe 2025 in Dublin, Ireland.
Talk: How to solve a Python mystery by Aivars KalvÄns
https://pretalx.evolutio.pt/djangocon-europe-2025/talk/PGQKUW/
Aivars KalvÄns presents practical techniques for debugging Linux and production applications without direct production access or a traditional debugger. He shows how `/proc` exposes process environments and open file descriptors, including deleted-but-still-open logs, and explains how `strace` reveals system calls, blocking operations, socket activity, and deadlocks across languages. Through incidents involving slow file storage, Kafka latency, dropped long-lived connections, and Kubernetes service meshes, he argues that debugging must examine the kernel, network path, timeouts, keep-alive settings, and underlying hardwareānot just application code.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Okay, hi. So my name is Ivers, and uh Shortly about me, I used to work for a software vendor writing software in C, so more than 18 years. And uh the software was used by tens of banks around the world. And uh Main thing was that we didn't have access to the client system, so every everything we had to debug the software was asking for specific log files or asking clients to execute specific commands. So that's one part how I got some of the skills. And now I work for another company, so luckily I do not have to ask for logs, but I still do not have access to production systems, so I can't uh use debugger as you normally would. And then I also do a lot of uh different projects and uh
Speaker 1: some uh cybersecurity stuff and this all comes handy But I have to start with a disclaimer. So I do not want you to get the idea that every project I touch is a complete disaster. However, uh If we get a single 500, then I will go and we'll have to investigate that. I will lose sleep and I will feel miserable while doing that. So uh and yeah, and it continues until I solve the problem. So let's start with uh some basic things. So there's this. If you know Brenton Gregg, he has these tiny books about observability of Linux.
Speaker 1: But uh he mostly comes at it from the performance tuning uh side. But however, a lot of tools of uh mentioned here can be used to understand what the system is actually doing. And I will uh Mention some of them as well. But first let's do something without any tools. So there's this process file system every Linux has. And we think of Unix so typical association is that everything is a file in Unix. However Uh processes as a files appeared first first in the version eight of Unix and it's around 1984 so it didn't start as everything is a file
Speaker 1: Uh so the most interesting uh piece of information there is the environment of the process. So it's uh Uh these numbers are the process ID and then slash environment. It's the file that contains all environment variables of report of the process. However, uh these are stored in the uh C style strings. So it means there's uh Terminating binary zero for each string. So what I typically do, I use the translation utility to replace binary zeros with new lines, and then I get access to the full uh environment and this is really really handy because if you are running your software in Kubernetes and you are using some let's say
Speaker 1: Helm uh home charts, uh some shell scripts, Python runners, then it's very hard to figure out w what exact variables will end up In your process. So this is where I go when I need to see how to log into the database. There is the password and typically the host and port of the database. Also, all your API keys are there So if I want to investigate something, I go there and uh find all the credentials without asking access to Vault or something else. The next very interesting folder is called FD. It stands for file descriptors, and under there you have all files that the process has opened. And since everything is a file, then
Speaker 1: Uh the first three are actually the uh console that I have connected to. There are some sockets, which are also files. There are some real files, there are some uh event polling uh kind of pseudophiles but they are there. So how can we use this uh folder? Uh the typical use case is restoring deleted log files. And for the purpose of demonstration, I have a simple uh Python script that is logging something to the ra uh file every second or so. And over the time it has produced uh something like nine close to ten gigabytes of uh logs. And a lot of people when they see a large log file
Speaker 1: or uh they are running out of space, they go and delete the logs. However, there's a big surprise. Uh the space is actually not reclaimed, it's still used So why is it happening like that? It's because uh uh historically Unix doesn't have a remove syscall It intentionally is called unlink and what unlink does it simply removes the file name from the directory. Uh the whole content stays there until the open file count drops to zero and then this everything is removed implicitly. So when you call unlink explicitly asking it to remove it, it will not remove it if the file is open by some process. And uh yeah, so this typically happens
Speaker 1: so when uh there were some problems setting up log log file rotations. Uh so the process is running but the log files are not being written to. So how do I find this? Is either you know the process ID and then you list this FD folder and you see that some files are have this uh deleted. Uh uh at the end of the name or I can grab for all deleted files uh in my system and then And then yeah, then I get the process ID, uh, which process is holding that file. So if my intention was to reclaim the disk space, then I just kill this process and the disk space will be reclaimed
Speaker 1: So most of the time I'm interested what's happening with the system. So there are two options. I can uh uh use tail minus f So just on this file descriptor, just as it was a regular file, and it will continue printing whatever the application is doing. The other option is to take a snapshot Of the whole content of the file. However, you can do like copy or move or something like that. You have to cat. So what cat does it's reads the whole content and then stores it somewhere else. Because this uh Number three is actually a symbolic link, so uh copying what just copy the symbolic link. And uh using similar approach I can find out which applications are using a specific file and keeping it open.
Speaker 1: So here I'm looking for Lua Snip and there are uh Fourth knee of him uh processes using it So that was like warm-up. The most interesting part is uh tracing. So as I said uh I couldn't use and I still can't use debugger in production so uh Uh we need to have a different solution and for that we use S trace. And uh so S trace uh What it does it actually traces all the system calls that the application is making. So uh it's calling the kernel There are many good things about it. So I don't care in what language your application is written.
Speaker 1: Is it Golang? Is it Rust? Is it Python? Is it PHP? I can use S trace on it and I can find out what your application is doing. Uh and somehow if I ask admins to run Strace on some application, they are more welcoming to do that instead of like uh installing a debugger and uh Typing some print comments in debugger. Another great thing is that it doesn't require root access, so if you want to trace your own processes, you can do it And yeah, if I have root access, I would probably go for BPF trace because it it's more it provides more information And the best, my favorite thing is that can it can also be installed without root access.
Speaker 1: If I do not have root access to the server, I can install it under my own user. And I have installed that even in a Kubernetes pod. So you just download two packages, it's one binary file, one library. You set up the library path and you run S-trace. So how do you run S-trace? S-rays, so this these are the typical use cases. I usually use something shorter, so minus F tells to follow all processes that uh will be started. And threads that will be started by the process minus TTT, it simply prints a very detailed timestamp. Minus O stores output to the file Minus S sets a bit higher limit on the strings that this uh
Speaker 1: tracing prints, and then you can attach it to an existing process that is still running, or you can start a new Linux command under that. Yeah, so that's the main idea. Uh so how does it look like from the running application? So here, so it prints you would say a lot of garbage, but once you get used to it, you can see that there's actually send two and it sends select distinct on loan unders underscore error And then it receives something that contains request ID and probably some value like 204 or I don't know. But it returns some value. You can kind of make out that it's doing SQL queries.
Speaker 1: Have a slower application, you would see that actually each sign consists of two parts. So there's the entering part which prints the uh input parameters to a function. So here is the pull function. Prints the input parameters and once the function returns it prints the result. And in the same way, uh when we try to receive data from socket 12, file descriptor 12. It waits in this uh part in the entering part and then it once it receives the data it prints the exiting part with all the data is received and also the length of the data. And uh so if you haven't done much
Speaker 1: uh system programming then this is uh Cheat sheet. So if you see E pol, P poll, select, P select, so it's basically doing socket multiplexing. So it's waiting on multiple sockets for some incoming data or incoming connections. And this is a fine state where you see your application hanging. If it you shouldn't see it hanging on the right or send to. So these should typically be very very fast. There are some conditions where it might block, but then the system is uh under very bad conditions. So if the application is hanging on the entering part of accept, this is also fine depending on how your application is implemented.
Speaker 1: So I would say I I don't see m many problems there. However, if it is hanging on the connect, entering part of connect, so it means that the application is not implementing timeouts. And I gave this talk uh three weeks ago and two weeks ago sentry fixed the bug. So uh when uh application was experiencing network issues, uh and it was trying to report uh exception to sentry, it could hang up forever because sentry didn't implement timeouts. And uh similar to connect, if uh the application is hanging receive, sorry, receive from read, it means it's waiting for data and again timeouts are not not implemented, then this is probably
Speaker 1: Either you forgot to set it up in the configuration parameters or application doesn't do it at all. Wait for it's typically waiting for termination of child processes and the next interesting one is few text. So this is basically waiting on a lock. And if you are interested more, then there's this fantastic site with Linux syscall table. So it has syscall numbers, names, and it links directly to manual pages of the syscalls so you can read more about what specifically your application is doing. And another interesting part of S Trace is uh printed Right when you start the application. So right when you start the S-trace. So here it says that the application actually has three threads, and then it prints
Speaker 1: state of each thread of application. And then it follows so the uh 298 it keeps working, but the Two other threads are actually blocked, so they are waiting on Futex. And if you look the first parameter of Futex is the address of Futex. So these are waiting on two different Vutex is mutexes and it means it has a deadlock. So if the address would be the same, it most likely they are just waiting for somebody to release the lock. In this case it's a deadlock And now the story time, I will have several of these. So we had application that was performing slowly, and uh
Speaker 1: typically the task that should be completed in like 10 minutes, they were not completed for an hour or more. So what we did connected to the server and the CPU load was okay for that application. It was kind of high for arbitrary application but for that one was fine, but they had high uh IOLOOD so I use the that time nmon to monitor that or if you so but that's kind of old tool There's IO stat that should come with every Linux. And if you pass minus X, then the last column will say how much of the IO capacity you are actually using. It will print percentage And in that case it was around 80.
Speaker 1: So we found out PID of one of the background tasks and run S trace on it. And what it was doing, it was uh stating, so stat is basically checking the uh file permissions, file size of specific file, and it was looking up a lot of lot of files in this slash storage folder. That gave us a clue because I forgot to tell that application log files looked looked fine. It was doing at the same rate it was doing it before. Uh what was the root cause was that we had this uh uh custom file field with some Uh extensions that it was looking for file in several folders, kind of in active folder and backup folder.
Speaker 1: And by accident the background task was simply loading more objects that need it, and therefore it caused more uh file checks. However, it shouldn't be that bad. So the question is why the IL load was so high because we are just checking if the file exists and what is the size. And just to save time, so uh we had new migrated to new servers with uh storage attached networks. And uh there's the So this command checks uh typical performance of the database because it's 8k blocks. So file systems work with 4k. I do not have numbers for the file system metrics, but I have preserved this kind of database metric. So 8K blocks
Speaker 1: and so the new ones were doing something like 27, 28 megabytes per second. And the old one Hetzner machine that we had used for years before that, it was pushing around 12 times more that more than that. And yeah, so that was it was really a hardware issue uh or virtualization issue. Uh so this got got uh some improvement since then but uh One thing to take away is that uh it might be also that the hardware is at fault, not not your software. So then networks uh Uh we have a lot of integrations with uh other companies and typically what happens is that we agree which IPs we will use and then both
Speaker 1: uh parties open. uh holes in their firewalls for those IPs and then they say okay you are good to go you start the application but it's not working And then you ask, oh, are you network admins, are you sure that you have correctly set it up? Yes, of course, it's a problem with your application. And then you as an application developer has to have to prove that uh It's something in a network and you have to do it without networking tools. And if you Google for that, Google will say that you can use ping and trace root to do that, but most admins typically cut off ICMP and UDP traffic, so those tools do not work. So what to do? So there's this wonderful tool cool called TCP Trace.
Speaker 1: So it tries to establish connections just like your application would. Uh but it uses this uh time to live to print a nice network I don't know, trace. So it connects to you and since it's it works with TCP it also requires a sp specific port. So this is a I tried to obfuscate some parts of it, but basically you run TCP trace root on some uh HTTP support, you get uh path uh that the packets took to reach that and then you show this to your admin and then they will tell oh why are the connections going now through incorrect IP? Well I don't know figure it out
Speaker 1: So that's a good way how to provide information to them and then they fix it. So story time. So we were building uh one of the first services using Kafka, and uh we had a setup where we consume a lot of CDC events. So uh Debezium was capturing uh changes from Postgres database and then we were uh enriching those messages and producing new events. And somehow it worked only at around 20 events per second. So some of people told that, oh, well, that's fine because Kafka can scale, and we can scale it up to region numbers we need. Yeah, for me it somehow seemed that it's a bit too little.
Speaker 1: So I got some time to go and investigate what's happening. So first I turned on the logs for the uh Kafka library. So it said that uh well it really took uh 43. 15 milliseconds to produce a message. Uh then I did a TCP trace route uh for this uh uh hosted Kafka and uh the kind of network times were around the single milliseconds so the network latency was not then not an issue. So and then I resorted to my favorite one which is S trace. And uh when I looked at S race, so I highlighted two timestamps, so it's when the poll
Speaker 1: so Kind of socket multiplexing call started and when it ended and so it means that it waited around uh forty-seven milliseconds for a response from Kafka. Uh so then I understood that okay, so it's not a problem with our application and uh I had a hunch, but I didn't expect that Kafka was that bad. So whenever you s you can't get past uh in your networking application past uh 40 milliseconds or push more than 25 events per second. So this is a magic number uh that is usually blamed on Nagel's algorithm So therefore all your sockets should set this TCP
Speaker 1: no delay that uh disables Nagel's algorithm. However, I've The real issue is the interaction with between Nagel's algorithm and delay deck, and I think that Nagel algorithm makes more sense. uh I would find a way to drop the delayed X, but yeah, that's me. Uh so we actually reported an issue to Kafka library and it's been around two years So I got a comment, uh seems nice. It's still not merged. So there's a non-default configuration parameter you have to set to disable this uh Nagel's algorithm. And there are some more links uh interesting to read.
Speaker 1: And we have another story. So this time it's a timeout issue. So uh the application was doing kind of some quant stuff, calculating stuff for And the application itself was written in Python and it was doing all the heavy work. There was a Scala client. That was calling the Python application, it was running in Kubernetes, and well uh a single HTTP request took a lot of time to complete depending on how much data you posted to it, then it the time changed. So the problem was that if you post one, two, three rows, then it takes like seconds, one minute, two minutes, three minutes, five minutes, and then
Speaker 1: Two plus hours and you get an error. So of course it worked in the development environment, it would it worked on the developer machines It didn't work in Kubernetes. However, on Kubernetes it worked with if you used curl to call the endpoint. So So whenever you have some networking issues, you you have to ask, what would curl do? Because this is the golden standard of network programming So if you look at the S-trace of it, first it says TCP no no delay. So perfect. Curl knows it. But then it says some keep alive and keep idle and keep interval So what are those? Those are uh keep alive settings. And why does the curl do that? Because Linux defaults are really, really bad.
Speaker 1: So in Linux defaults it says that it will send first keep alive packet after two hours of inactivity. And then it will keep doing it nine times with 75 second interval. That's too bad. Uh and you might say, oh well that's Scala stuff. Well no, Python is not better So one thing Python does better than even Kafka does, it says sets this TCP no delay, even in some crappy HTTP library, sorry. It's just HTTP library, it's not low latency, high performance, Kafka stuff. But it sets no delay, but it doesn't sit set the those keep alives. So uh and then again, I mean it's HTTP, you shouldn't be doing requests that take minutes over
Speaker 1: over it. But uh the thing is, we all have long-living connections Oh sorry, first slip. So what was happening that was uh some uh firewall and gateway was dropping connections after five minutes of inactivity. I don't know where where it Where it was, it was hosted Kafka. Sorry, hosted Kubernetes, no idea. But the client didn't receive uh fin packet, it means it didn't know that the connection was dropped The server didn't know that connection was dropped and it continued to work and tried to send response. And then the client set starts sending keep alive after two hours and then minutes later it understands that oh connection is timed out So we have long-lived connections elsewhere. So typically it's database connections that we try to keep alive.
Speaker 1: So there are settings, these are the Postgres settings. And uh please do set them because Linux ones are bad. And then uh so my favorite approach, but this is me I like the brute force so there's a library that you can preload and it overrides all socket creation and automatically sets this skip alive uh values to some reasonable ones. So with that you can solve it for Python, PHP, Scala, whatever. So this is what I use to fix the Scala issue. And once you think, okay, so now I'm good, but your application is running on Kubernetes, and the admins decide that service meshes are nice.
Speaker 1: service mesh layer four proxies are even nicer. So what happens is that uh so layer four means that uh it will uh intercept the real connection so you physically connect to this proxy and between application and proxy you have the keep alive and then the proxy creates a new connection to the database and it doesn't set the keep alive. So Uh yeah, so th this so we actually solved this by removing uh service mesh. Win win. Um yeah, so uh as a summary. So Uh the process file system contains a lot of goodies.
Speaker 1: I showed just a few of them. There are many, many more, but it's probably a bit too low level and uh Not for this talk. S Trace is a perfect debugger for any black box application. I do not care if your application is Written in Python, is it using uh clean architecture, clean code? Does it follow solid? Does it not? Is it complete spaghetti? Was it generated by AI? I can trace it with debugger. Then TCP trace route. So again a debugger for black box networking. Then remember to set timeouts for everything And remember this forty millisecond magic number. So it has been years and still there are companies posting uh Making posts every year that oh we solved our latency issues by setting TCP
Speaker 1: node delay to one. So just keep in mind, then set up keep alive explicitly for your long-lived connections. And then compare this. Your database server servers, cloud servers, against your laptop. So that will give you a good benchmark Sometimes I have written like SQL benchmarks to measure managed databases and the spoiler alert A lot of these are much worse than your laptop. So sometimes the issue is not with your application but with the hardware of the cloud providers and their setup. So that's it. Thank you.
Speaker 2: Uh my my question is how many years of experience do you have and how did you go to the Linux and Wall level stuff? Because it's usually people don't go there.
Speaker 1: I'm old and I have been working more than 20 years in IT. So and if you debug uh C<unk> applications you get pretty early into that Because in C<unk> if you make multi-treading wrong, it crashes. If you get it wrong in Python, so maybe once a year you will have an issue.
Speaker 3: Hey did you have any problems with uh t uh TOS or SSL uh connections as well for the keyboard stuff?
Speaker 1: Yes.
Speaker 4: Have you recommendations for tools for tracking memory, particularly um shared memory buffers?
Speaker 1: Oh shared memory as in system five Shen yeah?
Speaker 4: Yeah. So we've it's our uh any IPC memory. So We've server run out of memory, but it's not visible in processes, not VSS or any of the process tables. Don't see it in top, but the server runs out of memory.
Speaker 1: Yeah, it's uh still under process file system, so uh as well as the IPC uh message queue, so that's my BET topic IPC message queues. I like to use uh uh BPF trace, but it requires root access so uh then you can uh trace those calls inside the kernel. But yeah that's Uh I understand where I where are you coming from? It's a tricky tricky space. Because if you modify that memory you affect other processes. So yeah.
Speaker 5: Yeah, so um we currently have an issue with one of our servers and we really s we're really struggling to figure that out. And one thing is that we really have trouble figuring out when the issue actually starts happening. So it everything wants fine and then it does not. So uh I'm wondering is there any good solution to run S-trace in the background for some way to just figure out the time when something happens?
Speaker 1: You can you can do it, but uh depending on So there is a standard package in Linux called audit D where you can specify which system calls you want to intercept And it will store that to a specified log file. So that's probably uh uh kind of more I would say user-friendly so I have set it up like uh just because something happened bad during uh backup So I just needed to have like well becap process did these system calls around this this time so I can correlate uh things that happen. The other thing is uh well I would go for BPF trace if you have root access because it has
Speaker 1: More information but then yeah. Just to get the initial idea I would use audit T.
Speaker 5: Thank you.
Read `/proc/<pid>/environ` and translate its null-byte separators into newlines. This reveals the exact variables passed to the process, including database credentials, API keys, hosts, and ports.
Discussed at 2:27Deleting the directory entry with `unlink` does not remove the file contents while a process still has the file open. The space is reclaimed only after the process closes the file, so you can find the process through `/proc/<pid>/fd` and stop it if necessary.
Discussed at 4:45Use the deleted fileās descriptor under `/proc/<pid>/fd`: `tail -f` can follow new output, while `cat` copies the current contents elsewhere. Copying the symbolic link itself is not enough.
Discussed at 6:17Use `strace` to trace the applicationās system calls, regardless of whether it is written in Python, Go, Rust, PHP, or another language. It can trace your own processes without root access and can even be installed in a Kubernetes pod under your user account.
Discussed at 7:07Waiting in `poll` or `select` is usually normal socket multiplexing, while being stuck in `connect`, `read`, or `recvfrom` often indicates missing network or read timeouts. A thread waiting on a futex is waiting for a lock; threads blocked on different futex addresses can indicate a deadlock.
Discussed at 11:47Use `strace` to see what file operations the application is performing, then compare storage throughput on the affected server with an older or local machine using an I/O benchmark. In the example, the applicationās excessive file checks exposed storage that was roughly twelve times slower.
Discussed at 16:20Use `tcptraceroute`, which tests the path using TCP to a specific port, much like the application does. This works when ICMP and UDP are blocked and provides network administrators with evidence about the route and where it goes wrong.
Discussed at 17:37The delay can come from the interaction between Nagleās algorithm and delayed ACKs, rather than from the application or network path. Set `TCP_NODELAY` for the relevant sockets; in the Kafka case, the library exposed a non-default option to disable Nagleās algorithm.
Discussed at 21:02A firewall or gateway may drop an idle connection without notifying either endpoint, while Linux waits two hours before sending its first keepalive packet by default. Configure TCP keepalive explicitly for long-lived connections, and check that service-mesh proxies preserve those settings.
Discussed at 23:29Use Linuxās `auditd` to record selected system calls to a log file so they can be correlated with the time of the failure. If root access is available, `bpftrace` provides more information, but `auditd` is a practical starting point.
Discussed at 30:22Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025