Django migrations, friend or foe? Optimize your Django migrations for faster testing Denny Biasiolli

This video features Denny Biasiolli at DjangoCon US 2023 in Durham, North Carolina, USA.

Django migrations, friend or foe? Optimize your Django migrations for faster testing Denny Biasiolli
0:24:56
Published November 22, 2023
378 views

Django migrations are a great tool, but after years of changes in a project they can become very numerous, slowing down tests.
Is it possible to optimize them?

This talk was presented at: https://2023.djangocon.us/talks/django-migrations-friend-or-foe-optimize-your-django-migrations-for-faster-testing/

LINKS:
Follow Denny Biasiolli 👇
On Twitter: https://twitter.com/dennybiasiolli

Follow DjangCon US 👇
https://fosstodon.org/@djangocon
https://twitter.com/djangocon

Follow DEFNA 👇
https://www.defna.org/

Video production by the presenter and DjangoCon US 2023 volunteers.

Summary

Django migrations record model changes and provide commands such as `makemigrations`, `migrate`, `showmigrations`, and `sqlmigrate` for creating, applying, inspecting, and rolling back schema changes. Denny Biasiolli explains that many small migrations are convenient for development and atomic commits, but can make test database creation far slower than the tests themselves: in his example, setup took about 20 seconds while the tests took one second. `--keepdb` helps locally and `MIGRATE=False` skips migrations at the cost of extra setup work, while ordinary squashing mainly reduces the number of migration files and does not improve test database creation. To speed up tests, he demonstrates rebuilding an app’s migrations into a single initial-style migration while preserving the old files and `replaces` metadata, handling custom code, data migrations, fixtures, and circular dependencies as needed; this reduced the example’s total time from about 21 seconds to 6–7 seconds.

Key takeaways

  • Django migration commands create, apply, inspect, print SQL for, and roll back database schema changes.
  • A large migration history can dominate test time because creating the test database applies every migration before tests run.
  • `--keepdb` is useful locally, while disabling migrations trades migration setup for a separate database-creation step and is less effective in the example.
  • Standard `squashmigrations` reduces migration files but does not by itself make test database creation faster.
  • Recreating migrations as a single initial migration, while following a staged release process and accounting for custom code and dependencies, reduced the example’s setup time substantially.

Summarised automatically from the transcript.

Transcript

2,783 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:21

Hi there. Today we will talk about Django migrations, if they are friends or maybe not too much. But first, a couple of words about me. I'm Danny. I live in Italy and I'm working as front-end developer for fingerprint compliance services But I'm a full stock developer working with JavaScript, Python and a little bit of Go Today we will talk about migrations in Django. Migrations are a way to propagate changes to models into a database schema. And in Django, we have useful commands to Handle migrations. So for example, make migrations, migrate, show migrations, and SQL migrate.

1:11

Let's see them one by one. So for example, make migrations Creates new migrations for your apps. Its syntax can contain dash dash empty in order to create an empty migration And also you can customize the migration name for the file created and you can use this command just for a specific application if you want to pass the parameter So for example, if you want to create a model called Tweet with these fields, then you can save it in your models, in your main application.

1:56

And then running main migrations creates a new file in main migrations initial. pi containing the create model suite. Inside the file we can see the initial equal true because this is the first migration. Then we can see a list of dependencies in this case we depends on the out user model so that's the first dependency we need to use and then a list of operation For example in this case we have just a create model. This file is generated by Django automatically so we don't need to change it In this case.

2:44

Now, after creating the migration, we can update the database schema using migrate. The syntax is this one, so we can use a specific application label for the migration. or a specific migration name if we want to just update to a specific migration. For example, in our case we can round manage bi migrate. on a new project and all migration will be applied for example in admin out content types and so on We can also roll back migrations to the first one, also before the first one, using the zero parameter

3:31

as migration name. all migration will be undone. Then we can move to a specific migration like this. So specifying the application and the migration name Or if we just want to update to the last latest migration, we just need to run migrate admin in this case, and all admin migrations will be applied. But how does it work under the hood? Well, we can inspect the Django migration table in our database We can see it will contain four columns, ID, app, name and applied. And in this table we will see

4:17

in the app the application name, of course in name the migration name containing the number and the name specified by us or automatically added by a jango And in apply it we can see the time when the migration is applied to the database. There is a simple way to do this in order to just see if a specific migration has been applied or not, or also to see all migrations for a specific app We can use show migrations. In our case, if we want to see all migrations for our main application, we can launch this command

5:04

and the result is this one DX stands for okay this migration has already been applied to the database Then if we want to print the SQL statement for the name and migration, we can use SQL Migrate. as Django command specifying an application label and a migration name and the output will be this one for example for the First migration of main, we can see the create table, the alter table, and the creation of the index inside the migration. So what about we want to change a model?

5:50

For example our model extending the maximum length of the text of the tweet from 140 to 250 characters. Well, we just need to change the model in our files and then make a new migration Migration will be called with a nice default name, but we can customize it if we want. Then we can inspect the automatically created file containing the first initial migration in dependencies, and in operations we can see our alter field. If we want to see the underneath

6:35

SQL, this is the command to run, so we can see the ultra table inside the transaction. And then we can migrate our main in order to apply latest changes Now the problem is in the real world, what if we want to apply further changes? So for example, enabling tweet likes, adding for example a like model Enabling retweets, for example a new level field text and relate bit field Oh yeah, I forgot the related name for like tweet field, so I had to add it.

7:21

And keep in mind I'd like to during my development, I'd like to use atomic commits so for each single change I'd like to add a migration or a specific commit containing just this single change. So changes can be a lot, for example enabling followers and so on. Each one of them will result in a migration. In this case, like tweet related, alter -like tweet and follow. We can just show migrations and apply them. Let's make another real-world example. For example, we need to add a new shop

8:07

application. with for example a customer model and shipping details then we need to add this premium field to the customer model Then the business asks us to create a dedicated shipping address model in order to customize them and keep track of them. Then we need to migrate data to the new shipping addresses table with a new Django migration. and so on so removing customer shipping fields and again other changes in the future running on So increasing the length of shipping address, adding order model, adding create a depth to the order model

8:57

Order line, customer type, choice field in order to select between free and premium for customer. Then we need to migrate the customer type from the flag to the choice field. Then remove the is premium field. And then again, business asks for bronze, silver, gold, and platinum customer types, so we need to add them. Rename product quantity to quantity just for Clarification and so on, so adding product model. This will end up in a lot of migrations So yeah, I created this in the

9:44

custom repository I will share at the end of this presentation but yeah they are pretty simple to to use for development because we can roll back to a specific migration do the work and move on and again they are pretty easy. So we are happy, right? Well, not yet because well in talking about performances in tests, it's not so Nice and so easy. Well, as a disclaimer, uh timing may change from laptop to laptop, but in this case, timings

10:30

are Check it in my production machine, so keep in mind that. But for example, about performances, if we run tests in 20 applications like the one I created in as an example in shop. Running test will take about one second. So That's great, right? What is the problem? Why this presentation? Well, if we time our tests Yeah, it's true, running test takes just one second. The problem is in creating the test database, so before running tests.

11:17

a test database is created and this takes twenty seconds to run. Yeah, twenty seconds it's Not so much, but the problem is it's 20 times more than the running tests. So that's not a good thing, in my opinion. And also in our case running tests on GitHub Actions ended up in spending a lot of time creating the database more than running tests. So that was our problem and I tried to solve it in different ways, spending a lot of time And here's a quick recap. So first possible workaround in your local machine

12:05

can be keeping the database, the test database. So this dash dash keepdb preserve the test database between runs If the database, of course, does not exist, then it will first be created and migrations will also be applied in order to keep it up to date. So that's great if you add more migration they will be applied and then for subsequent runs of your tests The database is already there, so it's perfect. Well, there are a couple of pros and cons. For example, the pros are it saves 20 seconds for each test run after the first one, and that's great.

12:51

The problem is that it's not so easy to configure in CICD because you need to cache your database somewhere And that takes time too. So that wasn't our solution. Another solution, another workaround can be migrate equal false. By default is true. When you set this to false, migration won't run when creating the test database. This is similar to setting none as value in migration modules, but for all applications. So that's also a good solution. Single line change in your code base and it doesn't run migration during tests The problem is when launching your tests, it's like make

13:38

migrations and migrate before running tests. So in our test repository, this will add five seconds more running the entire testing suite. So that was yeah a nice solution with a single line change but that's that wasn't our Best solution. What about uh performance? So a quick recap, keep the B, nothing No time spent in creating database, but one second running tests. Migrate equal true, then 25 seconds, 5 seconds more. Nah, I don't like it. So

14:23

running through the Django documentation, I found out squash migrations and this will squash an existing set of migrations Into a single new one because I was thinking, well, if there are a lot of migrations, then we can just compress them into a single one and run it. So instead of creating For example, 26 transactions for our 26 migrations for each application, then what about creating just a single transaction containing everything? Well, okay, I a little bit time to spend and let's try it

15:09

So I tried to launch our squash migrations to shop, trying to squash migrations from the first one to the 26th And after a few seconds, this has been the output, so created new squash integration. Well uh there was a manual porting required because if you have some kind of custom command, custom SQL or custom Python code to run then that won't be migrated so you need to copy paste that into the squash

15:55

migration file I did that in the example repository you can see later on. And then you just need to specify which is the code to run for executing that. So, inspecting the migration file, please keep in mind this slide is important for later on. Inspecting the migration file, we can see that a new Value is there replaces with the list of application and migration name That this migration, this specific migration will replace. So reading through the documentation, the recommended process is squash

16:45

Migrations, keeping the old files, commit and release to production. Then wait until all systems are upgraded with the new release. And only then you can remove the old migration files, com it and do a second release After that you can transition the squash migration to a normal migration using this. So deleting all migration files it replaces. Updating all migrations that depend on the deleted migration to depend on the squashed migration instead. And last but not least, removing the replaces attribute in the squash migration.

17:31

Then you can commit release and everything will be perfect. Oh in Django, starting from Django 1. 1 You can use also manage PI migrate-prune in order to remove prune references to deleting migrations Only if you want of or if you think you want to reuse the name of a deleted migration in the future. So this will remove references in the Django Migrations table. Okay, now we have a single migration for every single application, of course.

18:18

Then let's test the performances after squashing. Hmm, okay. Well, a lot of time spent on this, but after squashing. Again, creating the test database takes 20 seconds. So that's not good, of course. So what's the point? Well, you can use squash migration only in order to move back from adding several Android migrations. To just a few. That's the single reason you need to use squash migration. So for example, you have a P branch you are working on with a lot of migrations and you don't want to

19:04

commit a lot of migration files Okay, then you can squash them, commit them and create your pull request with a single migration file. That's the only reason you can use squash migration. Yeah, I know, I know you wanted to speed up tests. Also me was like that so do you really want to speed up database creations in tests You need to recreate migrations from scratch and do a lot of manual tasks. So are you ready? Let's start. So first you need to annotate migrations for a specific app

19:51

like this. So show migrations for your application. store them somewhere in a file then you need to create a python list with this format so replaces of course do you remember and shop or application name and migration name in a long list After this, you need to move migrations in a temporary directory, so your current migration will be removed from your application. And to check that make sure that migrations are no longer there, you can run show migrations. Then

20:37

recreate the first migration from scratch using a different name than the old migrations so like this dash dash name init squashed And it will create a single file containing everything for that specific application. Then you can copy-paste the replace list we created. A moment ago in the migration file. You can restore all migration files. We moved into a temporary directory, we can move back in there. And then you need to ensure that all migrations are still there. So we show migrations, you need to see

21:23

everything from 1 to 26 in our case. Then you can launch the migration command in order to apply this in its squash. And then again back to post squash tos. So commit and release, upgrade old system, remove old migration files, recommit and do a second release, and so on. Then everything should be perfect. What could possibly go wrong? Well, if you have migrations providing initial data, then you need to create a new migration file for that. Or even better, please use fixtures. Look at the documentation for that. You can use them in test

22:08

and also application deploy, so they are perfect. Another problem could be circular dependencies because well there could be a couple of circular dependencies if you don't use a grain of salt in your development. So in order to resolve a circular dependency and break out one of the foreign key You need to of course remove the foreign key, create immigration, and move the dependency on the other app with it. Then move back the foreign key and create the migration.

22:54

So two migrations and you should resolve everything. Now it's time to test performances after recreating the migrations. And whoa, yeah, five seconds creating the test database because of this. first single migration file and of course one second creating the Running sorry running tests. So instead of 21 seconds we moved to six seconds. Yeah, I think that's great. So thank you for your time and here

23:39

it is your link to the sample repository you can see in the jugo settings migrate branch the change I did just for the single line of migrate equal false then in the other branch and this and in this request you can see squashing migration So changes for squashing the migrations and in the last branch recreating migrations Then you can see the last solution I found for moving from 21 seconds to 6-7. And that's great.

24:24

Thank you very much for your time. If you have questions, I'm here or you can contact me. everywhere you want. Here's my details. Thank you very much again and bye.

Questions this talk answers

What are the main Django commands for working with migrations?

`makemigrations` creates migration files, `migrate` applies or rolls them back, `showmigrations` displays migration status, and `sqlmigrate` prints the SQL for a migration. The talk demonstrates these commands for creating, inspecting, applying, and undoing schema changes.

Discussed at 0:21

How does Django’s keepdb option speed up tests?

`--keepdb` preserves the test database between test runs, applying migrations only when the database is first created or needs updating. It saves about 20 seconds on subsequent local runs, although caching a database for CI is more difficult.

Discussed at 12:05

What does Django’s migrate=False test setting do?

Setting `MIGRATE = False` prevents migrations from running while the test database is created, similar to setting migration modules to `None` for every app. It can save migration time, but the tests then need to create the schema another way, which took about five extra seconds in this example.

Discussed at 12:51

How do I safely squash Django migrations?

Run `squashmigrations`, manually port any custom SQL, Python, or commands into the generated migration, and keep the old files initially. Release the squashed migration, wait until all systems are upgraded, then remove the old files and update dependencies before a second release.

Discussed at 16:45

Does squashing Django migrations make tests run faster?

No. Squashing reduces many migration files to a smaller set, but it does not significantly reduce the time needed to create the test database; it is mainly useful for keeping a branch or pull request from containing hundreds of migrations.

Discussed at 18:18

How can I reduce Django test database creation time?

Recreate the app’s migrations from scratch as a single migration while preserving the old migration files and their replacement metadata, then follow Django’s staged release process. In the example, this reduced test-database creation from about 20 seconds to 5 seconds, cutting the full run from roughly 21 seconds to 6–7 seconds.

Discussed at 19:04

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Denny Biasiolli

More videos from DjangoCon US