Closing session
Published June 13, 2025
This video features David Wobrock at DjangoCon Europe 2021 in Online.
DjancoCon 2021 WorkShop | Migrations and understanding Django's relationship with its database
David Wobrock
Migrations are a very convenient aspect of the Django framework. They allow making changes to your models when needed, and impact the database schema iteratively in a smooth and integrated manner. No need to have a deep knowledge of SQL, be a database expert nor administrator - it just works. Or at least, most of the time.
The generated migration files reflect the model changes from one version to the next, and Django logically expects these migrations to be applied for the database connections to work.
The required synchronicity between code and database schema is the root of some issues one might encounter when using Django. We will dive into these issues during this workshop.
We will first explore in which cases one can run into these migration problems, and how they are intrinsically linked to this synchronicity. This will be done by creating a Django project and adding toy features to it, like any developer would do during the workday.
After defining the concept of backward incompatible migrations, we will also expose some example operations and why they can turn out dangerous.
The workshop will go about suggesting some existing solutions to these problems: we will manually fix such issues in development, but also explore how to prevent them from happening in a large-scale infrastructure with multiple servers.
Hopefully giving the attendees a better grasp of what is happening under the hood when something seems off with models and migrations.
Django uses the ORM for data operations and migrations for database schema changes, representing a graph of steps that brings the database into line with the model classes. The speaker demonstrates common migration problems: application-side defaults are not database defaults, divergent branches can require merge migrations, and code and database versions can become incompatible during deployment. To avoid downtime and rollback problems, he recommends identifying backward-incompatible operations with Django Migration Linter and applying changes progressively—for example, deprecating fields before dropping them, adding nullable fields before enforcing constraints, and using database defaults where appropriate. The transcript ends while he is demonstrating the django-add-default-value package, so the final solution is not covered.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Hi everyone, thanks for being here on what I call the migrations framework or in the long sentence migrations and understanding Django's relationship with its database. Sorry for the logo of my employer, don't worry, this won't be an advertisement talk. I doubt anyone would even be interested in the solution. But however, I think the slide templates look really neat. Just ignore the logo and let's keep going. So first of all, I'm David, briefly introduce myself. I'm just a little guy doing some Django, trying to roll up my sleeves from time to time and get my hands dirty. I'm French German. I live in Paris. I work at Botify, the purple company with the cool slide template. And I worked a little bit on tooling around Django migration. So I'm going to try and share around that today.
Speaker 1: So our plan for today is going to be in four steps. First we'll introduce a bit Django and its migrations, try to get up the basics and have this little toy project we'll use for this workshop We'll talk a bit about the theory behind it, then we'll try to add some features to the project. Like you would do, like any Django developer would do actually when you're working and you we may run into some issues and so we'll try to understand those uh formulate them why we have those issues and try to understand them uh and then finally of course we're going to suggest some solutions how to prevent them So I'm not 100% familiar with the remote workshop format. So I might be moving quite fast sometimes. Maybe too fast, don't hesitate to ask in Slack in room secondary if you have questions
Speaker 1: and that's not obvious. I'll try to keep an eye on it. There's also going to be some live coding, so I'm going to do all the changes in my terminal too. to illustrate it so um you can do all the changes on your side too it's always easier to write code yourself and execute command line commands yourself than just looking at it so feel free to do it but I'll be also doing it on my side. So the requirements, you might have seen them, maybe not. It's not mandatory. We'll have a little Django project. You can clone it, it's and this link. There's you can install the the libraries you'll need, which is mostly Django and um database uh driver for my SQL or
Speaker 1: PostgreSQL have a database running configure it and basically run migrate and run server to check if it works So a top toy project is about trees. Why? Just because they're pretty cool, it's sustainability, this ancient creatures. Well, I like trees. So the Django project is just a bare project with one app called Trees. There's one view, URL, and template, which are all linked and accessible under /forest. And there is a basic tree model with the with the first migration that creates a tree with a name. And that's all. So let's get it up and running. I already have a little terminal here where I have most. of the the project already set up.
Speaker 1: So I cloned the repo, created a virtual environment, all that. And I'll just run migrate to see if that works. Okay, there we are. It migrated. Now I'll run server to see if that works. And let's go here and we're going to go There and the server works. Okay, perfect. If we go into slash forest as we just discussed, we see that we have a very nice looking page with a lot of CSS skills here. My forest is composed of nothing and we can create a tree. So I create an oak tree and I have a tree. Okay, perfect. Very easy, simple toy project. So the very burst
Speaker 1: basics of Django migrations, you might know them. These are two commands, make migrations with migrate, that you probably know of by heart. I mean they're dealt with in the official Django tutorial in part two. So It's just after part one creating a project, part two, uh doing migration. So it's essential and rightly it's it becomes a daily part of your life as a Django developer. Let's take a step back and see what's a migration, what are they used for? So a bit of theory. So Django defines itself as an MTV framework, model template view, which is known by ordinary mortals as MVC, model view controller. And which interests us here is the model layer, which is here to interact with the database and abstract it from the rest of the application
Speaker 1: so that you don't have as a developer no need to you don't need to write the SQL statements to manipulate the database um the framework will do that for you and how does Django do that um well there are Python classes as we've just seen the tree class and that these these classes will represent the model layer, which pretty appropriately inherits from a class called model , which will handle this uh this model layer which is the database layout actually so these Python classes will be there to represent what the database should look like So a little quiz if if you will give you not like 15 seconds to uh um to write in the in Slack if you want.
Speaker 1: So What are the two types of operation that a framework, in this case Django, would need to abstract entirely the database? The two types of operations. What do we need here I'll wait maybe ten seconds, not not much longer. So The first type is CRUD operations, data manipulation. And there Django comes with an ORM, an object relational mapping. uh which is super powerful you already used it probably three objects also you f you're creating all these python objects um and they will be fully loaded uh um functional objects with everything that is filled.
Speaker 1: So here we have three we will have three objects with the name that is filled and that comes from the database. And you didn't have to know what database, how does it work, what do I do? And you can do all kind of data manipulation with the RM. The second type are schema manipulations. So that's always about metadata around the data itself. The tables, the columns, the types, the constraints. indices and whatever and these are handled through migrations. So why do we need these migrations? So um why What's required for this RRM to work correctly under the hood so that we can just use it without having to think of what it's actually doing
Speaker 1: I'll give the the answer straight away. It's that the database table should contain all the fields from the Python model class. So if we have a treatment name, we actually expect uh some table in the database with a column of names. So if if we put that the other way around, is that all changes to the Python model, these classes, should be passed on to the database schema. And that's what a migration is. It's the sequence of operation that will apply the changes to from uh on the database and the the the Python class onto the database and that will get the database from one state to the next to the one that
Speaker 1: Django is actually expecting. So when you take all migrations of a project, you have this graph of migration that will actually take an empty database to the final state That is mapping the model classes inside your code. So that's very short introduction of what's the migration framework and how does it work. And it's super useful. It's very iterative and that's why Django is a framework for uh for professionals with a deadline because your requirements always changing and so in your code and here you're you're able to modify your code in your classes easily uh without having to think about what does my my database schema look like actually
Speaker 1: what SQL do I have to write is it the PostgreSQL syntax which one is it you just have to change your Python class And that will be applied automatically. It's consistent in the sense that it's run in a migration. And we won't go into the details of your consistency model chosen in the database. In most cases everything will go fine by default. Also have a deterministic output, but that can be altered through the last point here that the migrations are very flexible. You can run any , you can modify your migrations manually, you can write raw SQL statements, you can execute any Python code too. So you can basically do whatever you want. So that's all cool, all great, but um let's let's dive into the the real stuff, like what's the smile
Speaker 1: workshop is about um so apart from uh forgetting to run make migrations migrate with hap which which happens to all of us all the time uh let's explore some more pitfalls So, to our trees project, a new feature is needed. Our users want to add uh want to also to store the height and the plant here of trees. Okay, well then Let's go into a terminal. I'll go into my editor and to my trees models. Here we have the actual model. I'll add a height which is a positive integer field. Perfect. It's going to be nullable because we might not know. Oops. We might not know the height
Speaker 1: of a tree we already have in our forest, and we'll have a plant here. And here the requirements are we don't want a null value, we always want a default that's going to be zero if we don't know when it was planted. Okay, good requirement. I'll save this. I'll make the migration. I'll stop the server for now. I'll make the migration. It created a migration the migration number two with the with both fields. I'll migrate that. And I'll run the server Okay, let's load our forest again. So our oak tree has no height and plenty of zero. Perfect. If I'm adding um Let's say a pine
Speaker 1: of height there's no unit but let's say 200 of whatever and then 2010 Works like a chart. Perfect. Our features up. We just implemented the new feature that was asked. Let's um well let's wrap that up into a nice pull request as we would do to get some code review. um and that our coworkers can have a look at it. So I'll check out a new new branch that I'll call feature. I'll add both fields, both modified files. I'll add a comet, add height and plenty to the tree model. Perfect. And I would push this, I would create a merge request, pull request, whatever, so that coworkers can look at it and say, hey
Speaker 1: Adding these two fields is fine. Perfect. Well that that that's that's sweet. Okay, let's let's go to the next task. Um oh our first bug. We already have a bug on our platform. That's that's unfortunate. What does uh this what does the task say? A user from Colorado tried adding the Rocky Mountain Bristle Combine, which is a 31 characters long name. Oh We could have seen that one coming, right? So by the way, this is what this uh pine looks like. And um okay, let's fix that. We should we we should have thought of that earlier, right? We're gonna increase the size of the name to 100. Let's say that will be enough for now. I'll go back to my terminal. Let's check out the main branch because we'll do the fix right
Speaker 1: on the main branch so that Well, we want the spug fix to be released as fast as possible. I'll reload my model file and I'll increase the max length to 100. Perfect. I'll save that. I'll make the migrations. It altered the field name on the tree model. Perfect. Now I'll migrate this. It worked as expected. I'm going to comment it right away because that will be easier for later But basically I would say okay my my bug fix is done we'll test it just after say okay that's fine let's commit this and say okay bug fix This is not a good commit name, but for the workshop it will be enough.
Speaker 1: Let's run the server and see if that worked. So here we are in forest. I'll reload this. The two columns are gone since they're in the feature, obviously. And now I'm gonna add the Rocky Mountain Bristol Compile Let's create that and it didn't work. Oh, we have a null value implant here of relations 33 that's violating a net null constraint. So this is an error that you might that you often receive in Django when you're developing features. Marcus Holterman discussed it two days ago in his works in his talk. about writing safe database migrations, it's a common case here that you might encounter. So what does our code look like?
Speaker 1: We have a name of length 100 What does the database table look like? So here I have a screenshot of the inspected database and we see four columns, the ID, which is auto-generated by Django. It's always adding an ID which is the primary key of the table, the name, which is of size 100 and cannot be null. Perfect. The height, which is an integer and is nullable, and the plant here also an integer which is not nullable So the error says yeah, okay, plenty cannot be null. It has a not mole constraint. That's true. Okay, that's This makes sense, but wait a minute, didn't we specify a default value to our uh to our plant here? We said it was zero. Well, time for a little
Speaker 1: We'll stop in our feature development. Let's learn about Django and its default values. So it was already touched on two days ago by Marcus, but let's it's Look at it again. So here's a new command in Django, which is called SQL Migrate. So when you run run a migration, when you migrate it, Django is generating SQL statements and executing them against your database with SQL migrates, you're seeing what it what are the SQL statements that are generated and executed. So here we have a little snapshot of the features branch, the migration number two, which is adding our both columns, our both fields. And we see that for the plant here it generated two statements. First of all, it added the column plant here, which is an integer with the default value zero.
Speaker 1: It's not null, and we check that it's positive. And then second, we drop the default value. So here Django is first adding the row and filling all existing rows with the default value, but then it's dropping it. Why is Django doing that? Well we discussed it just before that Django is very flexible about its default values too. There can be any Python function. And any Python code means it's not trivially expressed in SQL For instance, it could be calling a third-party API. How would you translate that into SQL? Even further, how would you translate that into different SQL dialects? I'm pretty sure PostgreSQL and SQL Lite
Speaker 1: do not have the same feature set for default values. So that's why the key takeaway that's written that's written in bold here is that the default value in Django in Django is handled on the application side, not by the database. Application side meaning it's the framework, it's Django, it's the J it's the Python code itself which is handling the default value. So it's Django that will see, hey, you're trying to insert a tree that doesn't have a plant here. It has a null value. But I know that I should I should substitute this null value with the zero. And then you do an insert statement with the zero value in the database is accepting that. So you might wonder here why
Speaker 1: why I mean how does Django fill the existing rows with a value if if it can be any Python function since here we still generated an SQL statement that will actually uh have a value and and and fill the existing rows. Well this here in this case we just have a default value that's a zero so it's pretty easy to interpret and to to substitute But if it would be a third-party API call, or maybe we want to fill all existing rows with random values, the default value would be a random function. Django would execute it just once to create this SQL statement. And then all new rows would actually execute the function, which means that all your existing rows would be filled with one and the same value
Speaker 1: even though you define it as a random integer. But then all new rows will really apply this function and have a real different default value. So the key takeaway here is obviously that it's Django that's handling the default value. So when I switch back to the bugfix branch, well Django there wasn't aware anymore of the plant here, so it didn't substitute the null value. The null value wasn't even in the insert statement. And it just failed. Okay, so how would we fix that? Here a bunch of different solutions with pros and cons. I won't go over them In general, you'll get the slides just after the talk. But some some of the solutions would be check out the feature
Speaker 1: and run the reverse migrations. That that sounds simple. You just come back and drop the columns. That has some drawbacks, which would be that you're dropping the data that's inside these columns But that could be fine. I could manually revert the database changes, just connect to my database, and if I know how Django works internally, I could drop the columns and edit some other tables so that everything's just as before That might be faster if I'm if I'm familiar with this. An easy way is to drop the entire data database and recreate it from scratch which would work just fine here since we don't have that much data and uh we just have two migrations but if you have a thousand migrations that might take some minutes Or you could make you can have a real default value as a database default, but we'll see that a bit later.
Speaker 1: Let's go with solution one here to fix our bug. Will we run the reverse migrations and undo adding these two columns? So let's go back to a terminal. I'll stop this one here. I'll go back to the feature branch. Check out feature. Here migrate Python Manage PY. migrate the trees to migration number one. So that unapplied the last migration. So now if I come back to my main branch where I have the bug fix and I run the server. I'll go back to my forest. Let's try again this Rocky Mountain Bristol Company. We get it We create it and now it works fine
Speaker 1: because we dropped both columns and we don't even know the plant G exists in the database. So that's happy. We'll now wrap this this comet up also in a bug branch and get it merged into production so it's released and our Colorado user is happy. So let's imagine we do that and now let's uh let's rebase our feature branch, right? We want to move we want to move forward and also fix that one since we have it won't work with this. So let's go back to our feature branch and rebase it against the main branch. So we'll do that real quick. We'll do a rebase main. There's a little conflict obviously because uh The name columns are not the same in the model.
Speaker 1: So here I reloaded the tree model and I want to keep my name with length 100, remove the name with length 25, and keep out to columns. So here we have and the box fix and our two new columns. I'll save this. I'll add the preest model and our rebase continue to end my inter my rebase. Okay perfect. Let's test again our feature. Let's migrate it again. So we add the two columns and An error that you might already have encountered quite often. Missing merge migrations. So we have a conflict in our migrations. Do you know this certainly? You worked um
Speaker 1: of a certain version of your code and a coworker or maybe just another feature you're working on is also of the same version. You both create a migration and when you merge them you have this conflict saying, hey, uh Django doesn't know where to go. We just we said it before quickly that migrations in Django are graph. Each migration might depend on one or multiple migrations So this creates a graph which Django will then discover and go onto the leaf nodes to know what is the state that represents all my model classes. in which I should be, in which the database should be. And here as we can see in this little in this little chart, right,
Speaker 1: Django cannot know should it alter the name length or should it create both columns, right? Both are valid states, but it cannot know which state is the final one that we expect. It's we as the developers, the engineers behind that that know, okay, it should be both actually in this case. So we have two solutions here. First of all, the one suggested by the command is running a make migrations merge, which is creating a little merge migration which is basically saying hey both of these migrations have to be applied and the the last valid state is this very last one where both already apply. This is pretty easy and straightforward and will straightforward and will work in nearly all cases.
Speaker 1: But it might create well the the graph is not as as nice even if you even more if you have a lot of developers working and you have a lot of um migrate merge migrations which will conflict with other migrations that can create a migration graph which can be quite a headache but Django is handling that, so that can be fine. Second solution would be to bump the last migration. That creates a pretty clean and straightforward migration history But you might have some rich issues on your database if you actually already applied these migrations and you you try to run again a migration doing the same operations This might fail. So it depends on your case.
Speaker 1: For example, you're running on you created the height and plant here here on multiple environments already, maybe different uh tests, sandboxes, staging environments, pre-production, then you might want to go with the merge migration because dumping the migration would need to either unapply the migration on all these environments so that it works or faking the new migration everywhere We'll go here with with bumping the migration. So we know exactly what's happening. It's just applied locally in our dev environment. So let's rename the migration. I'll go back to my editor. I'll go into my migrations. It's the number two
Speaker 1: auto migration. And I'll rename it to three. And inside the migration, I'm going to change its dependency to migration number two, which is called AutoTreeName, if I recall correctly. Let's save that. Let's go back here. Let's try to migrate again. And that worked. Okay, perfect So quick parentheses about how does Django know if it needs to apply a migration or not? It's quite simple in the end how it's implemented. It's a table in your database, which is called Django Migrations, which Django will store a row, an entry for each migration that it has applied. So here's a screenshot of
Speaker 1: just before migrating here this last this this migration number three of what it would look like So when migrating, Django is executing the content and adding a row here, telling itself, hey, I might I applied this migration successfully. There's a keyword argument, a par meeting but a parameter to migrate which is fake, which will actually prevent the migrate from executing the real operations but we'll just add the row here telling Django hey I applied that my this migration but you have to be aware okay it's It actually didn't run them. So that's in the case where maybe the columns already exist and you know that. So we just want to fake it and add a row here.
Speaker 1: You could also add it manually if you're a cowboy. Um okay, let's let's uh Let's keep moving. I mean let's go to our next feature. Project is never finished, right? So actually someone's telling us in the company the plant here is not precise enough. We actually need a complete plant date Yeah, that's that sounds fair. So I'm gonna edit my code, right? I'm gonna go here into my model again And I'm gonna edit the plant here, remove it. I don't want a plant here anymore. I want a planted on date, for example, which will be a date field with the default value, which is uh the today date. So date but today and I'm going to import
Speaker 1: from the famous date time built-in module date. Okay, I'll save that I'll go back here. I'll run a migration. It created a new migration number four, which is removing the plant here and adding the plant dut on. Perfect. I could migrate that. I can run the server. Oops. And check if that works correctly. Let's go back to forest, I reload it. So we have the height, which is none everywhere, even for the pine, because we created it and dropped the columns and recreated it. So it's null by default. And we have the planted on, which is June 4th.
Speaker 1: And yeah, that was fine. I could add another one, maybe a bird. And of height 150. That was yesterday. And let's create this already very big trick tree for for the plant date and okay that works. Now if we imagined the very same case as before where I would have to switch to another branch maybe the version just before we've planned here you might see it coming you'll get another error right you you um For the same reason that we just dropped the column, that's uh the plant here, and the previous code would try to add this column, but it doesn't exist, so that would fail. We won't do it here, you can do it yourself to to see by yourself, but um
Speaker 1: we have we're running into another issue. So how would we drop this column? We don't want it anymore. One possible solution to that, and uh I have heard many times, especially by DevOps or infrastructure teams, which are in charge of releasing, for example, the app your application and deploying it to production. They are often telling you, hey, your code stopped using plant here, this field. So the code is now not aware of it And let's just not drop it in the database since the code will just ignore this column. And in the case of a rollback, which is uh uh very good uh
Speaker 1: thinking uh from their side is if we roll back the column still exists and we don't need to restore a backup and just this one column from the backup which I don't know from when is the backup the last time, is it up to date? So let's just keep it. And since it's ignored, it should be fine, right? Well, what's the issue with this reasoning? It's not completely incorrect, but we are omitting the fact which we've seen just before. The default value is set by Django. But Django won't know about this column anymore, as they just said. Plant here is gone. Django will ignore it. So it will try to add a null value, but there is it's a not nullable field so keeping it would actually crash on insertion so that's uh a pretty bad bug in the sense that you won't detect
Speaker 1: it in development in development you'll say okay um I'm just running my migration, I'm dropping the column, everything works fine. But in production, you will try to do something that you thought is smart that will actually crash on every row insertion So let's let's take a step back and look at what's the issue with all these problems we we just uh looked at. So, okay, a lot of errors, they're bad and development. It's okay to just slow you down, but you can fix those. And to represent that with a little little little boxes here we have the code and database version one code and database version two and you see that code version one cannot work with the database v2
Speaker 1: since plan tier doesn't exist and So it has been dropped. And the code in version two has a column that cannot work in the database version one since the column doesn't exist yet. So the main idea is that Django in the migrations expect the synchronicity between the code and the database. That's very important to know. It's why Django is so easy to reason about and it eases the cognitive load of when you're working, of when you applied all your when you created and applied all your migrations Well, everything will just work fine and as expected. The real question is, is it possible to actually keep the synchronicity
Speaker 1: between code and database? Imagine a production environment where you're scaling horizontally. You can add servers, you just have REST API with no state that can that can you just pop um you know snapping your fingers and have 10 more instances to to handle your load. Well would you say that when I'm updating my code if do I update First the Django code or first do I apply my migrations? So this was also touched on by Marcus Tol talked days ago. If I update the code first, well that cannot work. I'm adding a new column and it doesn't exist in the database. That will just plainly crash. And if I update the database first, well, it dropped the column, so that also doesn't work.
Speaker 1: So I would have to update them at the very same time. And that the more databases and servers you have, the more difficult it will become. It will be very hard to have the deployment at the exact same moment. And so it will basically you don't have a choice but to generate downtime so that you can update them one after the other as fast as possible But downtime is not acceptable to anyone everyone, right? There's another case where uh where this will even if you if you would be able to deploy at the very same time There's another problem if we're in an environment where you have multiple versions. Let's imagine you're selling a stable version of your software
Speaker 1: uh v1 that only get bug fixes and you don't you don't add features to it um but it's targeting the same database as an agile version you also sell that This version has continuous deployment and it gets new shiny features every day, every week, every month for 2. 0. So If if now I had to apply a migration for a new shiny feature and that will break the database for the older version, for the stable version, that's pretty counterintuitive, especially for customers. The unstable version is breaking your stable version. That's the opposite of a stable version actually. So We will have a problem here in all cases.
Speaker 1: So what can we do? Oh we don't. Of course not. Of course we just looked at one specific type of migrations here, which are backward incompatible migrations meaning that when we apply them, the code uh prior to this migration will stop working. Of course a lot of backward compatible migration that we can add. For example, adding a nullable column to a table will just work fine in many cases. and that will be no issue. But here adding a not more column that won't have a default value or dropping a column or renaming a column, maybe changing its type. These are backward incompatible migrations and they are here in this case the source of our problems.
Speaker 1: So what can we do? Uh what can we do? Well first of all we accept our faith and say, okay, let's I'm just running a little application with one server, one database, they're maybe even on the same machine. I'm fine with a few seconds, maybe a few minutes of downtime. I'm happy to fix the environments manually and button but I like the straight-up or will be able to actually drop columns when I want them when I want to drop them or rename a field when I misnamed it in the first place. That's a fair trade-off. However, uh would be careful in the sense that even though you're okay with some downtime Maybe five minutes were it's possible.
Speaker 1: Um if you have backward incompatible migrations They're also very cumbersome to rollback, meaning it will double the rollback time in case you have an error. So first Deploying might take a minute or two. Maybe later it will take five, ten minutes. Then it will take ten minutes to see. Okay, uh we have a a bug in this release, we have to roll back. And then you need 10 minutes more to roll back the migration and to update all your servers. And however you during the update you will have mismatching versions, meaning that you have downtime So it can exacerbate in rollbacks even though you're fine with downtime. That's the big message. The solution we will be looking at here in this workshop is solution number two,
Speaker 1: avoiding backward incremental migrations. If you want to avoid downtime, you just don't do them basically. So that happens in two steps. First, well, you just don't do them. And how can you ensure not to do them is by using the appropriate tooling. It's hard if you're reviewing the code of a coworker who will add in migrations, maybe many migrations, to have a look at all of them one by one and to know which one will be fine, which one is not. So there is some tuning to help us with that. And the second thing, since you still want to drop a column at some point or as a not nil column, is to do them progressively. by doing only backward compatible migrations
Speaker 1: uh and be and ensure that you have no impact in actually having a hard constraint. So How to detect uh Django incompatible migrations? There's a little tool uh called the Django Migration Linter, which main goal is to detect backward incompatible migrations. So it's pretty easy install. We'll pip install it, add it to the installed apps, and run this management command link migrations. Let's do that. Let's jump right in. Let's stop the server. We'll pip install Django Migration Linter. Okay, that worked. Let's go into our editor. We'll go into the settings, which are in the migration workshop settings PY.
Speaker 1: Let's go into the installed apps and add the Django Migration Printer. Let's save that. Perfect. And that shall be it. Let's Python Manage PY Lint migrations and see what happens Okay, uh the linter went on all of our migrations, the four we have, and it checked for each of those if they are backward compatible or not. So let's look at it. First one. The initial one, creating the table. It's okay. Creating the table is fine in the sense that previous code couldn't possibly be aware of this model and trying to insert something. So that should be fine in this case.
Speaker 1: Migration number two, increase the size of our name from 25 to 100. So here the linter is saying there's an error. It's altering columns. But the little note is added. It could be backward compatible. You may ignore this migration. Why is it saying this? Here in our case, uh we went from trying in length of 25 to 100. That is actually backward compatible because the previous version of the code wouldn't try to add uh name that is longer than 25. So if we first increase the size, the previous code would still work. If we reduced the size, however, it wouldn't be backward compatible because the previous code would try to add a character of length 80, and we would reduce the size to 25, so the database would not accept
Speaker 1: uh a string of length 80 in this case. That's an example for strings. Same goes when you're changing the type of your um of your column, if it's going from a string to an integer or an integer to a string, one of those is backward compatible and the other one is not. Basically the linter is not smart enough to detect fine-grainedly if this change is okay or not. So in this case, we know as developers that this migration is fine for us. Migration number three. There's an error, not null constraint on columns. Yeah, we know we added the plant here, which is not nullable and has no default value. on a database so that's fine and number four we're dropping the column with
Speaker 1: plant here yes we are and uh not null constraint which is the planted on date field we added which is also not nullable So we we understand the output. So let's let's go let's jump right into it and try to fix all that, right? So case number one Let's try deleting the column. So generally you would do that in two phases. You would first make the field not used and accept no values so that you know it's it's uh it can be dropped safely in a later deployment. There is some tooling to do that and we'll look at it. It's called Django Deprecate Fields
Speaker 1: um which is there to do exactly that you will mark a field as deprecated instead of deleting it and this will um the package will also log when this field is actually accessed so that you know it will raise a deprecation warning so that you know that something is off and you shouldn't um and and that you shouldn't drop it just just yet So let's try and rewrite the last migration to see if we can if we can avoid dropping the column using this little package. I'll go back to the terminal. I'll start by unapplying the last migration since we already dropped it. We dropped the applied gear card here. Let's add it again.
Speaker 1: So I'll migrate back to the version three. So we unapply the migration number four here. I'll remove it. I'll remove trees, migrations, oh, four. And we're gonna edit our model. Let's go back to our model file. And here we now add again our plant here since we don't want to delete it. We just want to uh to to mark it as deprecated uh since deleting would would generate downtime so it was a positive integer field uh with a default zero And we're going to import from Django deprecate field, which we didn't install yet. We'll install just after, deprecate field.
Speaker 1: And the way it's used is that you encapsulate encapsulate your field. In this function. So it will be marked and overridden by the library. So let's save that. Let's pip install the Django deprecate fields library We'll install that one. And now we should be good. It's made the migration where we don't drop the field this time. You see that it added the planted on field and it altered the plant here. So okay, let's migrate. Nothing else too, right? Let's move on. Let's migrate this. That worked. And we can now check maybe through the SQL migrate command what it actually did.
Speaker 1: So trees or migration number four. And we see that the plant here we actually dropped the not null value And additionally, this tool will now log so you can detect easily when this field is used, even though it shouldn't. Now if we link the migrations again, we'll see that migration number four is not dropping the column anymore. Cool, we're one step closer to making the linter happy. Let's go to another case which is um which is often a bit more interesting. Uh it's to add this non this uh add a not-nullah column So you would generally do that in the three-phase process where you would add a nullable field. Then you would ensure that your code always sets this value
Speaker 1: so that you know as a developer that it cannot be null and the code before the migration will always specify this value. And then when you are 100% sure you make it null Or uh you can use a database default when that is possible. So we said if it's a third-party API call that inserts default value. You might want to do the first strategy and keep this default value on the application side, but you can propagate if you know what you're doing, you can actually have this database default on the day, this this default value on the database Two possibilities. Either we add it manually or we add it with some tooling. We'll try both here in this case. Let's try adding this
Speaker 1: manually on or planted on date field here. I'll go back into the migration. So let me open up the editor. I'll go and edit I'll go and edit the migrations file Migration number four. And in this case here, we'll add a custom SQL statement ourselves. So that okay, I'm running SQL and here I want I want to, I'll copy paste it since this is some live coding. So here's the PostgreSQL statement to add the default value, which is now. So that's the the the current date. so that we actually have the default value on the database.
Speaker 1: So I'll save this and now I have the default value. If I migrated this We can lint the migrations to see if uh if if the linter detected that correctly. We see that there's a warning now for migration number four uh which says that the run SQL data migration is not reversible. Oh good catch Linthe thanks. Indeed, if I try to undo this migration uh it would crash on the run SQL. So we could define a no -op since just after that we should drop the field, but let's do that cleanly and actually do a drop default uh to reverse this operation. And here we go. I'll save that. And
Speaker 1: now we have A fully run SQL operation, that should work fine. And if we run the linter again, migration number four is all set. Let's try to get all our migrations working. So at least backward compatible in regards to the linter. So migration number four added this plenty here with default value zero Here we'll use something else to try to add the default value. We're gonna pip install a package called Django add default value. Django add default value I'm going to go back to our migration, this time migration number three. And after the plant here , we will actually add a default a
Speaker 1: new operation that will add the default value in the migration. So I'll add a the import that goes with it. So from the from our library we just installed at Django default value. I'll import add default value which is a Django migration operation, custom one, that is here too add a default value to to our um uh to the migration why why would that be more useful than doing it in manually well Before we had an alt-the-colum statement, which might be specific syntax for SQ for Postgres QL, you know, and the now function uh is maybe specific, but Django can handle multiple uh database vendors
Speaker 1: if you multiple databases at the same time so uh this is very specific to one and if you wanted to change then you would have to rewrite them potentially Whereas this package will generate the correct syntax for you depending on what database you're using So this is just defined as the migrations above. I'll just copy what's there. So we add a default value to the model that's called tree. to the column that's called chance here with the value well we define zero. We'll save that. We'll lint the migrations again And see that this time migration number three is now okay. We added the database default value. So it's not backward
Speaker 1: incompatible anymore. And now since we also know that migration number two, there's written error, but we know it's not an error, we 're aware of what we're doing, we also added this one. And as the lintel suggests We're gonna um ignore it. So we'll add an another custom operation from the linter itself. From my Django Migration Linter, we'll import ignore migration And we'll just add that here as a first operation. It's a no-op. It will generate no SQL. It will just tell the linter, hey, please don't lint me. I know what I'm doing. in evaluant migrations one last time now they're all valid and one is ignored
Speaker 1: and uh well with this knowledge we we now know right that our All our migrations are backwards compatible. We have the appropriate tooling to detect it so that we don't have to think about it every time we write migration. So it's automated. And automated is often a good thing. And we're going to avoid with this now the annoying cases during development or even during deployment in production. Basically avoid downtime. So the idea would be to make this, let's go back to this like this linter happy. And that will prevent this downtime just by automating the process of detecting backward and compatible migrations. For this workshop I I on purpose didn't use SQL Lite
Speaker 1: because all the versions of Django and SQLite are actually generating a different type of operations. You could try it yourself running the STL migrate on the migration number two we had with the database SQLite, which is already defined in the settings And you won't see the one liner you expected. For SQLite, the way it was handled is that it would create a new table with the new schema you're waiting. The new states that the model is expecting. It would copy all the data inside of it, drop the old table and rename the new one So that's much harder for the linter to detect if the change is backward compatible or not. So that's why for this workshop we use Postgres
Speaker 1: QL and MySQL has the same cases in here for uh for this workshop but that's the reason we didn't use SQL light. Okay, uh nearly out of time. That's perfect. Let's sum up what we discussed and and seen. So migrations are super cool and useful They have some pitfalls, but if you're aware that we need the synchronicity between code and database and that it's hard to achieve during production deployments because of backward incompatible migrations. Well, you have some techniques to to detect them and to avoid them so that everything goes smoothly. So it's all about your environment, right? Your constraints. If you know that you have to to avoid downtime
Speaker 1: or um or you have multiple versions running in parallel well if you know these answers to these questions you know how to avoid these migrations books now I hope you learned something. Thanks a lot for listening to me. And I'll be available for face-to-face and questions on Slack or on whatever other support. Thanks a lot.
Speaker 2: You put so much effort into this into this backwards compatibility compatibility because uh you have different versions on the same Data running.
Speaker 1: Yeah, at some point uh I worked in an environment where we had this constraint. Did it this very example where we're selling a staple version of the same uh SAS software, but also uh uh Yeah, and version where where we added our our features and sprints and everything and they both targeted the same database since it was shared data about some yeah just business data that had to be known by both. And we didn't want to duplicate the databases at that time. So we we were always breaking the stable version because of these incompati with incompatibility.
Speaker 2: Yeah, I understand. Okay. So the then then I get this point, yeah. Um
Speaker 1: But it can also happen if you just have not one I mean if you have many servers in production this also happens because you will have a mismatch when you're the when you're deploying the database and your code it won't be the exact same time so you still have downtime at this point if you have backward incompatibility migrations So it's not the only case where this actually makes sense. If that's clear.
Speaker 2: If you can ensure that you um as in the other talk, the other guy, I forgot his name, uh was mentioning that Yeah, maybe uh that um all those um operations that delete um uh attributes um you can shift in another release. You can maybe handle this, but this is not the case in your situation.
Speaker 1: Well it would be making the change in multiple releases basically, you know, ensuring you're not using it. And then it's uh Marcus discussed it very well saying I'm the engineer behind this. I know what I'm doing. I know this column is not new not used. I can delete it. And I'm doing this Well, one release at a time, knowing what you're doing, that's that's the best possibility here.
Speaker 2: Yeah, yeah. Understood. Cool. Thanks a lot.
Speaker 1: Thank you Any other question? Or something else you want to talk about?
Speaker 3: Hi, I I wanted to ask uh when the project becomes very large, uh is there a way you can just Remove all the migrations and just generate the new ones? Have you ever tried that?
Speaker 1: Do you mean uh squashing all your migrations into one? Or do you mean specifically for the linter I I showed in the the workshop?
Speaker 3: No, not the linter. Actually uh it's beyond squash. After one squash, you can't even do another squash. It says you have to undo the previous squash and then do the squash, I believe. So I have a project which has uh squashes and it has hundreds of migrations and um what I want to do is for the migrations to start over. So uh have you ever tried anything like that? Any experience with that? Um
Speaker 1: completely triangling start audio scratch.
Speaker 4: I believe my belly
Speaker 1: I know uh uh someone has been has started to work on something similar of uh trying to remove old migrations and starting from new migrations Um that I don't know where his project is. I think it's on and go on Gitter, but it's largely abandoned However, for your squash migrations, you should be able to squash again a squash migration. If when you squash them, it's explicitly marked as hey this migration is a fusion of a merge of all these ones and once you apply the squash migration and Django is aware that it exists and it's substituting all the others Uh you should delete the old ones and the squash migration should be marked as you should remove this dependency
Speaker 1: and then Django isn't aware anymore of hey this is a squash migration so it's just a common migration as any other one. And at this point you should be able to squash it again and to redo the same operation.
Speaker 3: Okay, so I have to remove the dependency and then it should work.
Speaker 1: Yeah. I don't remember exactly the attribute name that's that's in the squash migration But if you remove that and and this thing saying hey I'm depending on other migrations then that's uh that should work I I believe
Speaker 3: Okay, sounds good. Uh I read that to restart from scratch what you have to do is delete all the migrations, uh, create the new ones. Then run a fake migration that you are at the same level so that the database thinks you have applied all of them and then kind of do it. I've never never got to try it
Speaker 1: Uh that would work, but you have to be careful about your uh data migrations, your custom run as SQL or your run Python code, since if you delete it and you make migrations they won't be there. Maybe they're adding some um default data you always want or some indis indexes that you always want in your database. So just be careful with that since squash migration will keep them and just removing, deleting the migration file will just will it'll block.
Speaker 3: Yeah, yeah. Okay. If there's no other questions I have some more. Oh
Speaker 1: thanks.
Speaker 3: Okay. Uh so I I wanna ask another thing. Uh many times I have to do some changes in the database like I have to uh when I'm doing a migration and I've added a column and that's a computed column, I have to compute the values In that case, what I do is I create an empty migration and run the command either as a Django code or as a SQL. To uh do that as a part of the migration so that it's it goes on. Uh is that a good way or do you recommend uh not doing it that way?
Speaker 1: Marcus discussed on this two days ago and he had a good point. It's the migration process is often often blocking for your release. So if you're doing a data migration, computing a value on a very large table that needs to be locked to compute them, you You might have a migration that will take a very long time to run, which might be blocking your release or locking a table that's critical or not being able to acquire lot uh the lock So in the migration you're you often you often have some critical path which will be acquired. Whereas if you do um what suggests the mark is uh management command Well, you can run that at any
Speaker 1: time you want and uh maybe do in batches the update of your computed column. So it depends basically on the size you expect and the the The execution time of your migration I'd say.
Speaker 3: Okay. Yeah, I I it's a it's a balance I g guess, but uh sometimes you can't do the migration uh command because um migration command because uh your database state depends on it. So if you start the application without data in that column, uh it might fail as well. So you
Speaker 1: that should be a migration I guess if it's a requirement to have
Speaker 3: Yeah.
Speaker 1: Migration sounds sounds like the best option here to me at least.
Speaker 3: Okay.
Speaker 1: Any more questions or remarks?
Speaker 3: Uh no, I I really like the linting tool that you told about. Uh the add default, that's also good. I've never tried these. We never really focus on these things, I guess. Um migrations are kind of like a stepchild, uh they just exist. Um
Speaker 1: Well you don't use them until you have to, until you run into the issues and ask yourself like, hey, what did I do wrong? I just added the column. Why is everything crashing? That's the point where you say, okay, let's let's think a bit deeper about it
Speaker 3: Yeah, yeah. Um now that's all I think. Uh I did have one problem earlier, but that's probably not applicable here. That was uh using Django tenant schemas where you have to run migration across a lot of schemas. Um and if something fails, then uh you're in the middle where half of the tenants have are done and half of them are not done. Um it's kind of a really messy situation. Um So
Speaker 1: I'm not an expert on the tenants, so I'm not sure I would be able to help here.
Speaker 3: Yeah, yeah, I I understand. I understand. Well thank you so much for this session. I really loved it. Thank you.
Speaker 1: Thanks, cool. That's nice feedback. Someone else wants to ask a question or As a remark where it's useful or not?
Speaker 2: I definitely want to look more into that because uh currently it's not an issue. But as we are also providing RPs which have to be uh then downwards compatible, I think we will run into the same uh situation. And um pretty good pointers. Thanks a lot. Yeah.
Speaker 1: Cool Okay, well th thanks for this feedback. That's nice. Okay. If there's not any any more question, then uh I'll leave the Jitsi. Thanks a lot for listening to me and yeah, have a good day.
Migrations handle database schema changes—such as tables, columns, types, constraints, and indexes—while Django’s ORM handles data manipulation. They translate changes to Python model classes into operations that move the database from one schema state to the next.
Discussed at 6:16Django applies defaults on the application side because a default can be arbitrary Python code and may not translate consistently across database engines. During a migration it may use the default to populate existing rows, but it normally removes the database-level default afterward.
Discussed at 15:32The default is supplied by Django only when Django knows about the field. If code from an older branch no longer includes that field, Django omits it from the INSERT and the database rejects the missing value when the column is non-nullable.
Discussed at 18:00Django migrations form a graph, and parallel migration branches can leave multiple valid leaf nodes. You can create a merge migration with `makemigrations --merge`, or, when the migrations have only been applied locally, renumber or otherwise adjust the migration history to produce a clean sequence.
Discussed at 22:17It records one row for each successfully applied migration in the `django_migrations` table. Using `migrate --fake` skips the database operations and records the migration as applied, so it should only be used when the schema already matches.
Discussed at 24:57Adding a non-nullable column without a usable database default, dropping or renaming a column, and some column-type changes can break code that is still running. Adding a nullable column is generally backward-compatible, and increasing a string length can be safe when older code still writes shorter values.
Discussed at 34:15Avoid backward-incompatible migrations and make changes progressively. For example, add a nullable field first, update code so it always writes the field, and only later enforce non-nullability; similarly, deprecate and stop using a column before removing it.
Discussed at 36:34The Django Migration Linter checks migrations for operations that may break older application code. It can flag changes such as non-null constraints, dropped columns, and potentially unsafe alterations, although developers may need to judge whether some flagged changes—such as increasing a maximum string length—are actually safe.
Discussed at 37:19First stop using the field and make it nullable or mark it as deprecated, then remove the database column in a later deployment once no running code accesses it. The `django-deprecate-fields` package can mark the field and log accesses so you can verify that it is safe to remove.
Discussed at 41:17Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025