Run your tests in hundreds of different environments fast. I mean really fast.

This video features Anton Pirker at DjangoCon Europe 2022 in Porto, Portugal.

Run your tests in hundreds of different environments fast. I mean really fast.
0:16:55
Published October 17, 2022
995 views

Run your tests in hundreds of different environments fast. I mean really fast. by Anton Pirker

As a Python library author, how do you ensure interoperability with every Python (and Django) version? In this talk, learn how to run your test suite in hundreds of different environments as fast as possible using Pytest, Tox, and GitHub actions.

Summary

Anton Pirker explains how Sentry’s open-source SDK test suite, covering hundreds of combinations of Python, web-framework, and framework versions, was reduced from roughly 40 minutes to under seven. The main improvement came from running Tox environments in parallel across GitHub Actions runners and CPU cores, followed by removing unused services and configuration and moving Tox’s working directories to a RAM disk. He argues for iterative, time-boxed refactoring: make one small improvement, measure it, keep what works, and aim for better rather than perfect.

Key takeaways

  • Tox can run environments in parallel, but GitHub Actions runner limits determine how much concurrency is actually available.
  • Generating separate workflows and combining runner-level and CPU-level parallelism produced the largest speedup.
  • Removing unused Node, Redis, and unnecessary PostgreSQL services saved additional time and simplified the CI configuration.
  • A RAM disk accelerated creation of Tox virtual environments by keeping them in memory rather than on disk.
  • The speaker recommends time-boxing performance work and iterating through small, measurable changes instead of attempting a perfect redesign.
  • The project’s next goals were faster runs, potentially under two minutes, and overdue Django 4 support.

Summarised automatically from the transcript.

Transcript

2,953 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Man, so many beautiful people. It's nice. So my name is Anton. I'm I'm the guy from the Napkins. You've maybe seen. I know it's it's weird, but yeah. And I want to talk about how to run tests in hundreds of different environments really fast and the really fast in the title is like the key point so keep this in mind for later. But first uh what I have to maintain. So we have a open source library called century SDK And we have around four and fifty tests in our test suite. And we support around twenty Python web frameworks And also support very old Python 2. 7 and then 3. 5 until 3. 10. And also one of the those uh 20 web frameworks is of course Django, and we support Django 1.

0:45

8 and up So key uh 4-0, uh 4-0 and 4-1 is not quite there yet. I'm sorry for the delay, uh, but we're working on it, but I'm busy doing other stuff. Um Yeah, so all in all we have seven Python versions to support, times twenty uh Python web frameworks, times two to nine uh versions of each of those web frameworks. Which means we run our test suite in over four hundred environments. And this uh is pretty big, yeah. And how do we run this? Our Our stack for testing is like a basic I think that the default stack for everyone it's like PyTest for running the tests. We have FLAC 8, Black and MyPy for um

1:30

type checking and linting Then we use Tox to run our test suite in the different environments. So TOX is the the main part. It creates virtual environments for each of the Python versions and each of the framework versions and then runs the PyTest test suite in it. Then uh we use make for running tests locally, uh make it easier to run our tests locally. And as a CI provider we use GitHub Actions where we run all the tests. on each uh pull of a release branch or on each uh merge into uh the main branch and also like on each pull request. So when I was uh submitting this talk in the end of May, this is how our test suite looked like in GitHub Actions. And let me zoom into

2:16

the important part. This was the duration of our test suite. And that's not really fast and it's not even modularly fast, that's very, very slow. But why did I then submit the talk? Because like it's very very very slow. So my thinking was this Um if I submit a talk and it gets accepted, I can represent Sentry and Sentry likes this. So that's good. But if I uh come up with thirty-eight uh minutes of test duration deck, it's uh it's embarrassing So and I don't want to embarrass myself in front of you all. And also like I don't want to embarrass Sentry. So the thing was, maybe if I submit a talk and it gets accepted I need to make time to fix the tests. And also like it forces my manager to give me time to fix the tests.

3:11

Like it was between thirty eight and forty two minutes. So and this is very annoying if you like on each release and everyone you have to wait forever And also the one thing is also like if then the test gets accepted, I can fix the test suite. We have faster tests. It's a win-win for everyone that contributes to the to the project because just every developer's life gets easier with a faster test suite. Um when I told this to my manager, or when my manager Vladimir found out, he coined the term for it. It's conference-driven development, basically. And I like the term a lot, so like that's that's it for now So I gonna talk to you in this talk about how I my my journey from 40 minutes test three to faster and how fast I get and what I did So let's do some refactoring.

3:56

And by the way, uh all the images in my slides are done by the stable diffusion AI. It's like a it's a Python tool where you Type in random stuff like here Fernando Pessoa using a computer and it paints stuff. And it's like amazing and I spent way too much time playing around with this, but it's just too amazing to not do And also like Fernando Piazoir is an is a Portuguese writer, an amazing writer, you should check him out. So first uh Circle of refactoring. Basically everything is always circles. So first you have to find ideas how to improve my current situation. Then you pick one of the ideas and you implement it and then you measure uh has my situation now improved? If not, throw it away. If yes, you can go back to finding new ideas or adapt the ideas you found in the last iteration

4:46

And if you l look at this, it's not only like the circle of refactoring basically, it's also good advice for life in general. So it's also kind of the circle of life. So yeah. My first idea was uh I should run the tests in parallel probably. Because I I found out that when I run my tests locally it only uses one CPU and that's just Stupid idea. So that's the the the basic start I had. On the horizontal line is the time. So we had like this one GitHub runner. And the GitHub runway is like the virtual machine that GitHub Actions uh styles up and runs your workflow in. And in this, we started talks, and Talks then created a virtual environment for each of our four hundred um uh test

5:31

suites and ran uh each of our four hundred environments and run the test suite in it one by one by one. So that's the slowest version you can have basically. So I thought maybe TOX can run stuff in parallel. And it can. It has a dash-parallel auto, then it uses uh the number of CPU cores you have to run the environments in parallel This already greatly improved the speed. It was like from 40 to 25 minutes, but 25 is still yeah, not enough. And this was also like when I changed this, our test suit, some of the tests started to fail because if you run tests in sequentially or in parallel, it's different. But I I implemented this, uh we released it and I went on a vacation. So this is also

6:16

something you should not do. Uh because some of our tests uh start up a uh a server on a port And if then the same server is started twice on the same port, the tests crash. So my uh my colleague Neil, who's also maybe here somewhere, fixed it. Thanks Neil. So uh here the limiting factor is the CPU course. In GitHub runners you only have two CPUs. So that's all you get. So I thought if I do not get more CPUs, maybe I get more runners. So the next idea oh sorry was to get one runner for each of the environments we have. And for this I needed to change our uh GitHub Actions configuration file so that we had one YAML file for defining all those

7:03

uh how to run our test suite. And because I'm a lazy uh engineer, I wrote a script to like split it up. So on the top left you see or on the top right? Uh your left. You see the uh the TOX configuration file where if where we define all the Python languages we use and all the frameworks and all the framework versions. And I wrote a script that parses this file and then creates for each Framework we have uh one YAML file uh for configuring GitHub Actions. And so again, this is what I thought I would get But in reality you also have like not unlimited number of runners you can start. And we have a GitHub enterprise account, like a big account, and this means you get one hundred and eighty concurrent

7:52

uh runs for the whole organization. So it's not me in century, there are a lot of other bright people that run stuff in CI. So it was more like this. I had five to ten or maybe fifteen concurrent runners. Uh depending on how much all the other people at Century do. So this did not uh improve the time that much So I thought uh maybe I should do something like having the base best of both worlds, uh parallelizing with talks on the CPU level and on the runners. So I changed my script that creates the the YAML files To have it like this. Now um I start one runner for Django and in it I start talks with uh all the Django versions. So the test between all the Django versions.

8:38

And this brought the time down quite a bit. So The result of this all was it saved around twenty to thirty minutes of duration of our test suite. So the next idea then uh was I can clean up our YAML file. Because when I did the script that creates those YAML files, I noticed some stuff in our YAML file that's just not there anymore, it not not used anymore. Like we have a setup node in there and I I don't know why we need node to run our Python tests, so I deleted it. I also noticed that We spin up two services in all of our test runs. It's Redis and Postgres. And then I did some digging and found that in our code we do not use this Redis service at all because we have now something that's called fake Redis.

9:24

It's like an in-memory Redis for testing. So I deleted Redis. And Postgres was only used by the Django tests. None of the other frameworks used it, so I changed again my script that generates the the config files to only start a Postgres instance on the Django. Framework. With the result again two minutes saved. The next thing was uh I should cache uh I used a cache of GitHub Actions. So In our GitHub Actions, uh Tox is trading this virtual environment. It's like a directory where it installs uh the dependencies for the tests from PyPy. And between runs I wanted to cache this directory so it's not installed over and over again. And there's something called Actions

10:10

Cache in GitHub Actions that lets you do this. But as our dependencies are set up, it could have happened that uh we installed the dependencies, then they were put in cache, and if meanwhile one of the dependencies updated Uh our cache would not be invalidated. So we run our tests against the old version of the dependency and not the new one. So long story short, I did not implement it because it could have uh Um yeah, it didn't work. So but I kept the idea for later. Um the next idea was then if I cannot use the cache, maybe I can use RAM drives. And RAM drops are basically you have a directory in your file system, but it's not sitting on the disk, it's sitting in memory. And um memory is still fifty times faster than disks, also

10:55

SSDs. And I was thinking so maybe I don't know I was worried that GitHub actions do not allow me to to create RUM dives. But it turns out you can do them. On the top left you see this uh I created directory with this mkdir and then the sudo mount uh command. This just assigns a certain number of megabytes in memory to this directory. In this case 180 uh megabytes. And on the t uh bottom right, uh In our TOXINEI, I just tell TOX to use this directory as work dear and tempdeer, and our tox creates the virtual environments not on disk, but in memory. Which again s saved two minutes. Sorry. Next idea was a local PyPy

11:41

server. So still the the virtual environments are created in memory now, but I still have to download all my dependencies from PyPy and then install them. So we have an internal uh PyPy server at Sentry. And Anthony Sutil, he joined a couple of months ago. You maybe know him. He's like he streams a lot of Python content and coding on YouTube and Twitch. If you don't know him, Anthony writes code, you can check it out. It's really cool stuff And he set up this internal PyPy server for speeding up our main monolith repos CI, which is based on Django and other things. On this internal Piper server, a lot of packages are there, but all the dependencies I needed, like all those ancient Django versions and stuff, are not there.

12:26

So I could have added them, but for now it's I did not implement it, but I kept the idea for later. Time's up. Oh time's up. Because conference-driven development has a hard time limit. And my time limit was a couple of days ago when uh JungleCon wrote me an email and said, Hey, JungleCon is around the corner, uh let's send us your slides and we can check them for code of conduct. And I was like I don't have any slides yet. So maybe I stop doing tests refactoring, but start doing the slides. And there's a sad programmer. So, as a recap, the stuff I considered. So, first was the first idea was like running the tests in parallel, and this took a couple of iterations to get right.

13:15

It was also the main part of the work. But saved the the most time. Then I cleaned up the YAML files, which is a easy thing to do and it's like saved also a couple of minutes The GitHub Actions cache was not able, but I saved it for later, and the RUM drive saved some minutes, and then the PyPy server is also something for later. So as a total results, the number you want to see is now our tests are now running under seven minutes, and most of the frameworks run in one or two minutes, but only three or four of them. Django included, around six minutes around. And this is also like Django is our biggest integration and we have the most tests for that. So yeah. Also like splitting up the test suite in in all those different uh like workflows in GitHub actions make it

14:01

way easier to find the failing tests. Because before it was like a humongous log of And finding the one test that failed was really really annoying. So this is also like a great result. And I did a conference talk to you all. participate now in so I think it's a it's a success all in all and also Cristiano Ronaldo is with me. Yeah so what is next? We uh so I have a a milestone called Better Test Suite in our project, you can check it out. And all the ideas I had in this talk I had I put issues in this milestone and the ideas I did not do are still in this in this milestone so you can check it out and also if you have like ideas on how to do this even faster and stuff just create an issue

14:46

put it in the milestone Or maybe I have to put it in the Meizen, I don't know. Uh but just put it there. Being under two minutes should be easily possible if I had had more time. But first we will uh focus on the Chango 4 support because that's long overdue. Uh as a conclusion, uh one thing if you do a like a big refactoring task like this or any big thing, uh is like you don't have to make it good You just have to make it better than it is now. And I think that's the main thing you should take away from this talk. It's like it doesn't have to be perfect, it just has to be better than now And also like iterate. Don't do a million things at

15:32

the same time if you do like a big task like this. But just pick small ideas, implement them, test them. If they work, keep them, if they're not working, throw them away. And the time boxing thing. Uh do time boxing. Conference-driven development is like you put yourself a little bit under stress uh stress and because you have a fixed time. Like a real deadline that you need to have something to prepare. And also like because when I did this I had to do my day to day work also. So it there's this small um Daily time boxing where you cannot uh restrict the amount of

16:17

work you do. Performance improvement will take you, but you can restrict the amount of time. You can say I will do three hours of performance improvements today and this every day. And this is in combination with iteration, it just makes makes uh the project better step by step. And I'm a big fan of this baby steps approach thing. Yeah, I think that's all I have from for me for today. Thank you much for listening.

Questions this talk answers

How do I run tox test environments in parallel?

Use Tox’s `--parallel auto` option, which runs environments concurrently using the available CPU cores. This reduced the suite from about 40 minutes to 25 minutes, although parallel execution required fixing tests that conflicted over the same port.

Discussed at 5:31

How can I run hundreds of CI test environments faster with GitHub Actions?

Split the test matrix into GitHub Actions workflows, but group related environments on each runner and let Tox parallelize within the runner. This hybrid approach avoids relying on an unlimited number of runners and cut roughly 20–30 minutes from the total duration.

Discussed at 8:38

What else can speed up Python tests in GitHub Actions besides parallelization?

Remove unused setup and services, start PostgreSQL only for Django tests, and place Tox’s working and temporary directories on a RAM drive. These cleanups saved about two minutes, and the RAM drive saved another two; dependency caching and an internal PyPI server were considered but deferred.

Discussed at 8:44

How fast did the test suite become after the optimizations?

It went from roughly 40 minutes to under seven minutes. Most frameworks finish in one or two minutes, while the larger Django integration tests take around six minutes.

Discussed at 13:15

How should I approach a large test-suite performance refactoring?

Make one small improvement at a time, measure whether it helped, keep successful changes, and discard or revise unsuccessful ones. The goal is not perfection immediately—only to make the system better step by step, using time-boxing to keep the work bounded.

Discussed at 14:46

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon Europe