Creating an Inclusive Django Community with Kenya Phelps
Published July 15, 2026
This video features Chris Wilcox at DjangoCon US 2019 in San Diego, California, USA.
DjangoCon 2019 - The blameless post mortem: how embracing failure makes us better by Chris Wilcox
In today’s world of developing services we tend to move fast and with that comes mistakes. This talk will discuss using post-mortems to turn incidents into opportunities for improvement, instead of just an opportunity to assign blame.
This talk was presented at: https://2019.djangocon.us/talks/the-blameless-post-mortem-how-embracing/
LINKS:
Follow Chris Wilcox 👇
On Twitter: https://twitter.com/chriswilcox47
Official homepage: http://crwilcox.com
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Intro music: "This Is How We Quirk It" by Avocado Junkie.
Video production by Confreaks TV.
Captions by White Coat Captioning.
Software systems and their failure modes become more complex as they evolve, so failure should be treated as a source of learning rather than a reason to assign blame. Drawing on transportation investigations and medical morbidity-and-mortality conferences, Chris Wilcox argues that psychological safety, curiosity, and independent facilitation help people share what happened, while prevention and concrete follow-up actions matter more than simply identifying a root cause. He demonstrates a postmortem structure covering impact, detection, resolution, timeline, what went well or poorly, what was lucky, and action items, using an accidental deletion of Google Cloud Python documentation as an example.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
So I wanted to start off on sort of a positive note since most of this talk can get pretty negative. Software is pretty incredible. In the 20 years that I've worked with technology I think it's safe to say that it's evolved. We've come a long way from where we once started. In the nineteen nineties, when I was originally um starting my journey into using a computer, The internet was a small thing. I would have never imagined the sort of speeds we could get now or the size of it. We've also come from a place where websites have come from a place where they were a very simple, small site, single purpose. It's having multiple services interlinked with complicated auth, complicated security problems, multi-regional deployments, and many more things.
And it turns out when your software gets more amazing, so do your bugs. That's part of the evolution. We don't get to have one without the other. So it's reasonable to expect as we continue our failure modes are gonna keep evolving. They're gonna keep getting more complicated and harder to deal with. And we're going to fail. So I was thinking, what if we saw this as an opportunity? What if instead of seeing this as a reason to run away from computers and say this is all a big mistake, we should have all just been gardeners and janitors. and never talked about these computers, maybe we could learn from these bugs. Maybe we could take these negative experiences and use them for positive gain. This is what postmortems give us. They give us a way to learn, to grow as a team, to deliver more things for our customers, to be more successful.
We can do this through failure. But there's an interesting thing. This only works if everyone on your team feels safe to fail. It only works if they feel open to share when they make a mistake so we can all learn from it. And then we can prepare better for future failures. So before I get too far into the talk, I just kind of wanted to take a second to talk about myself. I spent most of my time in my career working for large companies. I started working on. NET websites, but then I moved on to working on IDEs and compilers, cloud services, and building cloud platforms. Today I spend most of my time working on the client libraries for Google Cloud for seven different languages, Python being one of my personal focuses.
I think it's interesting to reflect on why we call it a bug. We have these these failures in our code, these problems. But we didn't just choose the word bug. The issue is at a certain point in time, they were literal bugs. We could see them, we could spot them, they caused problems in relays. And as far as I know, this is the first computer science postmortem. This is an example of a moth taped to a sheet of paper describing timestamped what went wrong and what the problem was. I kind of wish it was this easy to show the problems that our software systems have. I've never been able to tape it to a sheet of paper. But here it is. So our systems have evolved. The failures have gotten harder, but we still do the same processes.
The problem now is they aren't moths. They're cosmic rays, service interruptions, complicated failure between your two microservices and some off-service to use. So maybe it's not going to be single pointed, but we can still use this process. I also think what's interesting is nowhere in this is their blame. I mean, we don't blame the moth for falling into the relay. It happened. We can set up things so the moth can't get in the relay again, but blaming something isn't going to move us It's not going to prevent future failure. So we can look at our own past, but we can also look at other industries. There are a few other industries that I want to talk about that find ways to handle these complex scenarios in their environment.
And I think we can use these to look back at software. And for the industries I'm going to pick, they deal with some pretty serious subjects. I think at one point or another everyone says the phrase, yeah, it's life or death, but for a lot of industries that's true. That's what they deal with day in and day out. They deal with people's lives and the outcomes of their situations directly affect that. But they all have the same thing in common that I think tech does, and that we can learn from our mistakes and we can improve our systems. And if we can find ways to shift from blame to investigating systems, the reasons things go wrong, and put in place preventions, we can grow. The thing is, a lot of cultures and organizations
are gonna focus on blaming individuals. I would say the culture we exist in outside of tech defaults us to this. We want to look at someone and say it 's their fault. They did it. But if you really reflect on that, that didn't solve your problem. And at least for me personally, it doesn't really make me feel better. It makes someone else feel worse. It doesn't move you forward. So why are we doing it at all? So let's look at some other industries. I want to start by talking about the NTSB. The NTSB is a US organization. That monitors rails, highways, aviation, pipelines, and hazardous material transportation. They exist to investigate incidents, accidents, failure modes During transportation that affect
people. Like I said, a lot of their failures are pretty serious, right? Airplanes, when they crash, uh have some pretty disastrous outcomes. So there's an interesting aspect of the NTSB. They aren't prosecutors. They cannot use the evidence collected in their investigation to blame someone or hold them legally accountable They can prevent this in the future. They can take actions, change processes, change regulations. But they cannot use this incident to hold you at fault. We have other organizations. So in the US we have the FAA, we have the FBI. They're going to handle that part. But the NTSB is solely concerned with how do we avoid this?
This terrible thing happened. How do we not do it again? In that environment, blame has no place. So as I said, they've provided a lot of safety recommendations to this, about 13,000 of them to date. And they do this with a few techniques. Without blame, people are open to share their faults. And then we see from people's mistakes, even if they're inside the process, That's the faults, right? That's what caused it. There was something that went wrong. And we'll talk about an example from TNTSB in a bit. But if you blame them, they're not going to share that. They're gonna sit quietly in the corner and wait for someone else to say anything at all that can maybe explain this other than what they did. But in an environment without blame, people share openly and you can
Can move forward and you can start to figure out how a lot of times incidents aren't single-pointed at all. There's a few things that had to work together. And we have no hope of learning from that and preventing in the future unless we understand all of those aspects. So here's an NTSB report. They start fairly boring. I would say this to me is very recognizable with design documents that I've used in engineering. It describes a very factual situation. Something happened. In this case, they say it was an accident. Says if there are injuries. In this case, there weren't any injuries. And they say what did it cost? Ultimately, what was the cost of this event? So you don't really gather too much from this, you get some weather. And then they tend to include things like pictures.
And for me, this makes this very, very easy. Okay, so there's an accident, there's a collision. That boat appears to have a hole in it. Uh that's probably the problem, right? So how did we get here? How did we get to this this place? In this particular accident, they found that certain processes weren't followed, but they also found that the processes that exist didn't necessarily cover the situation. They were in. And so the flexibilities in the existing processes and recommendations allowed people to have this failure happen. So at the end of the report is where these recommendations come in. In this case, one of the things that wasn't happening was there was no one on lookout on a ship. And it turns out if you have two uh vehicles on an open plane and eventually they're going to intersect if no one does
anything to stop that. And so these boats ran into each other. And they're very lucky that no one was injured, but they treated it with the same seriousness as if someone was. And from this we're able to devise new situations for how you ought to communicate and how you ought to look out. I think this is probably the most important thing that we can learn from the NTSB. There are plenty of postmortems I have seen where we go through the work to find out what went wrong, but we often don't think about preventative measures. This is likely the most important thing. The rest of it that describes what happened, it's informative, but it's not actionable. The other industry I wanted to talk about was MM conferences.
So MM doesn't literally mean the Mars and Murray's candies that we We've all probably eaten entirely too many of. I mean morbidity and mortality. So it's a slight bit more serious. So the healthcare industry. Remember, I said at the beginning, some of these people deal with life and death quite literally. So the thing about morphedity and mortality conferences is they have very similar goals to the NTSB. The words change a little bit. They care about patient outcomes, they care about educating practitioners so they can do a better job with patient care. But this almost always is used not for your sort of routine failures. But your more complex ones, where possibly multiple physicians or facilities or organizations are involved when things start to get more complicated and you can't simply uh
describe it as a very single point. So I have a photo on my slide. I don't know about all y'all there, but this looks pretty adversarial to me. This is from a television show from the perspective of the presenter. And if people are looking at me like that, I'm not going to be very open to share. I'm not I'm gonna feel uncomfortable sharing my faults, that's for sure. I'm not about to tell you I did anything wrong because I think you're waiting for that. That's that's what you're looking for. It is a dramatic portrayal. It's from a television show. And at a certain point in time, this is what MM conferences looked like. It looked like senior physicians, hospital management sitting with you and telling telling you, you know, to
Profess everything you've ever done wrong so that way they can put the fault on you and it wasn't the hospital's fault, and that's it. But we don't learn from that, and people don't share in that. They they tend to they tend to get quiet. I'm hoping that it wouldn't surprise anyone that this isn't common practice anymore. The medicine industry learned that that dramatic portrayal, which was more common, wasn't producing the sort of result they wanted. They were hoping to educate people. They were hoping to inform physicians, organizations, hospitals on better ways to practice medicine that resulted in better outcomes. You weren't going to get that with an adversarial panel, but you might get that with a lecture hall. So they changed the setting.
You have now taken a situation that I think you would say was full of animus, uh, and you've replaced it with curiosity. This simple change of setting with the exact same thing can change everything about how people view this. There's another thing I thought was interesting when reading about MMs, is they have a lot of conversations about who should be present and who should be able to talk when present. This this related very much to me in tech. I've had similar conversations conversations about this, usually this takes the turn of should management be in the room? Should people who are not ICs, who are not developers be in the room for a post-mortem Because whether we like it or not, there's a power dynamic there. And so if your manager is in the room, you might feel less open to admit you've done something wrong.
And it turns out within uh within medicine, they're basically even split into two. So the stat I read uh from the the NCBI, uh the N the NAH, it's a subset Set of the NIH. Said about 50% of people think they shouldn't have management at all. So it's it's kind of comforting to me to learn that even across industries we have similar feelings. But I read about another thing they had. They used a concept called a chairperson during their postmortems. The idea of this was that the person presenting on the particular incident. wouldn't lead the meeting. They would worry about presenting the postmortem. And someone else would facilitate the room. They would control who spoke. And so if you allowed management to be present, for instance, you could make sure that they weren't able to
able to turn it into a very blameful environment. And so uh we use we use similar things like this uh in other areas within tech, but I've never used that for a postmortem. Um and I I personally think this is something that I would like to use in in my organization. after all my reading. There's another thing that I found that was interesting about medicine. Does anyone recognize this, what this is? So this is uh what's called a fishbone diagram or an Ishikawa diagram. It's a method used to tease out what caused an adverse effect. I've used this in two separate organizations. Sometimes it's surrounded by the fish, and it sort of makes it clearer what this is. They use this in MMs.
They use the exact same process. The words change. Right? We don't use these particular words. We don't talk about people and patients or equipment or environment or maybe a medicinal organization. But we talk about our auth system. or user actions or infrastructure actions the same way. What I learned as I researched is that for the most part, we haven't really invented any of this. But no one tells you that, right? No one tells you that this has mostly been adopted from other places. But we don't keep reflecting on those other places. We took it at a point in time and moved along, and they've evolved and we've evolved. but we never go back to see how we could learn from each other. So
what can we learn from these industries? I think we can learn to try to Treat each mistake as an opportunity. I think we can learn to adapt some of their processes and respect the opportunity we have in a failure to find out more about our systems. And I think honestly the the most important thing is to learn to take prevention more seriously than the investigation itself. I think in a postmortem it's really hard to sort of turn that around because we spend so much time worried about root cause analysis, and then we spend so little time on preventing it from happening again. This is the fun thing about about postmortems, right? They create a record in the history that we can use to learn from. But I don't really think we're learning from that history yet.
So more specifically, with all this and all this research, how might we conduct a post mortem? Postmortems ultimately only have, I would say, three components. We worry about why something happens. We think about how could it have been even worse than this? And we'll get into why that is later. This is the learning from a failure. And how can we make sure this never happens again? Because the worst scenario is that we have something happen, we investigate it, and we repeat it. history anyway. That last one needs to take precedent over new work. And that that is sort of going to be a review Repeated point. If you don't take the actions seriously to prevent things, you're just going to keep doing this.
So here's an example incident. This follows very closely a process I use today in my organizations. I've made it a little more general because we have some drill downs that are more specific to our product, and you probably will too. We start super high level. We just start with a summary of what went wrong. So here, the incident was that the documentation for the Google Cloud Python library was unavailable for users. Doesn't really tell you what went wrong, but it gives you a rough idea of what you're reading. This is useful mostly because you're going to generate a lot of these. You're gonna forget when things happen, and you need a quick way to find them. them. This tends to work pretty well for that. It's also useful for management leadership. In a big organization, I work at a rather large company, so it's handy that uh people above me can have a very high-level idea of what's going on and where are we having
Issues. The next thing you're going to talk about is what made it happen. Not necessarily what all happened, uh, but but how did this get triggered? And so after investigation, we found that as part of our cleanup process for our repositories , an engineer with write access to that development repository managed Delete the GitHub Pages branch altogether from the repository, which causes us to lose all of our GitHub pages documentation. Okay, well that's that's pretty bad. Um okay, next we're gonna worry about user impact. How long was it? How much did it affect them? In this case, for about two hours, we didn't have documentation. This means every user who tried to go to our documentation suddenly got an uh
oh 404, oh no, you can't do that. That's that's not great. It could have been worse. I mean no one likely lost money from this. For instance. So this is this is not a terrible incident. It's a little embarrassing, though. For the team as a professional organization to do something like this, it doesn't look very good. So how did we notice that our documentation went down? In this instance, we had a report from an external customer. And this sort of adds to the, maybe it's a little more embarrassing, right? Uh typically in here you'll find that this is detected by some bot you have or infrastructure that monitors your system or a dashboard alert. In this case, that's not what happened. We had a user that said, hey, I'm trying to use the documentation, it seems unavailable.
What's going on? And between the time that was reported, it took us 30 minutes to notice that. That's probably not great, but all in all, two and a half hours, maybe that's all right. So how did we resolve it? For this issue it's it's pretty straightforward. We republished the GitHub pages branch and the world kept turning and it was fine That's about where management's probably going to stop reading these reports. Beyond this, it tends to get very uh technical, it tends to get a lot more brainstormy. But above that line is what you're going to show to management. and help them understand what's going on with your system. It tells you how big the problem was, it tells you how you fixed the problem. It gives an idea of how long it took to fix the problem. And then they can look at things like resourcing or prioritization and Help you from that end.
But after this, we're going to get into the more devy side of things. What was the exact timeline of what happened? Not rough hours or anything. This is where it gets really interesting. So you start to see things like this is at 9 45 in the morning. I work at a tech company, so a good majority of the team probably wasn't even here yet, or they're still finishing other breakfast or getting coffee, right? And so that was deleted. Um okay, so some some some person came early to the office and deleted the branch. Um and then an issue was filed like 40 minutes later. And then, you know, 24 minutes later the team responded. And okay, they republished the branch 10 minutes later, uh, but 50 minutes after that, it was serving docs. So there's probably something interesting there that we didn't talk about before that from an engineering side we're going to probably want to dig into
why did it take 50 minutes after publishing for our docs to go live? I think this is uh an interesting area, what went well, what went poorly, and what got lucky of a post-mortem process. I haven't seen this used everywhere, uh, but for me this is sort of the fun part. You look at what happened and you start to think, is that was that really how it was supposed to work? Did we just kind of luck out? Could that have gone worse? You know, how did this happen? Like what thing could we add to this that would make it worse? you start brainstorming about non-existent failure. So in this case, we were notified pretty f pretty quickly. Our customer is able to get a hold of us on GitHub. Issues, but by the same note , we didn't know about it very fast, right? It still took us 30 minutes-ish
to realize it was there. The real thing that we noticed that went poorly is that you're able to delete the branch at all. Um I mean GitHub has controls for that, right? You can protect branches. And it turns out for this particular repository, we never enabled that because no one had ever tried to uh delete the GitHub A. Branch before, and so no one ever noticed that it wasn't set. And so there's a lot we can look into this. There's also a lot more notes about that 50 minutes I talked about that we found for our organization and the size of our GitHub repo. That we use for documentation, it takes entirely too long for us to figure out if we've failed to publish something correct and there's not a lot of logging we have access to. And so probably going to want to do something to harden that up in the future. And where did we get lucky?
Well, it turns out a member of the team had a copy of that branch. So we didn't have to regenerate all the documentation or any of the scaffolding. We just published it from their dev box and then did the minor update that would have otherwise happened. Not a big deal. It turns out that this was hugely helpful because without this, I think it probably would have taken us another two hours to get this online. From these things, from these final actions, from the incident, we can develop the most important part, the action items. So action items are going to fall into a category of five things. We're going to either have to investigate the incident further. That's a valid action item. Is that when you get done with the post-mortem you say we don't have enough information? This is gonna take take someone to take down time as a dev. And so we're going to have to have them look into that.
The other is that you're going to repair damage. You're gonna detect future incidents, you're going to mitigate future incidents, you're going to prevent future incidents altogether. So in this case, we had a few things. We're gonna turn on branch protection. That's easy. Not gonna let that simple of a thing happen again. But there's some other things that could still happen, even with that on. So we're gonna need to figure out how to better debug our GitHub page. And that's not a thing where we're going to settle in a post mortem meeting. That's a thing where someone's going to have to have an action item on their calendar to go take care of. And then we're going to also look and say, maybe GitHub Pages isn't for us. A lot of people do this successfully, and a lot of devs on our team, myself included, use GitHub pages. for smaller projects or on websites, it works great.
But when you start to have gigabytes of information, it seems to have a bit of a problem. I think it's important to mention again that we have to fail. Uh if we have any hope of innovating, we're gonna fail, we're gonna make mistakes. In my head, uh, this is what I hear when people say fail fast. I don't like fail fast personally. I think it implies that we intend to fail. That's not the intention. And that that's the other part is individual act with good intention. You don't have a team member whose goal today is to break your service. That's not why they came into work. It's an accident, and we can learn from the accident.
It's also important to remember that blame isn't a deterrent unless your goal is to deter people on your team from doing their job. That's the only way you don't fail, is if you just don't do anything at all. If you do that, you won't fail. It will work. The problem is that's sort of against the other main goal we have, which is to use usually accomplish new things, make money, innovate for our customers. And punishing a team member doesn't help you. It doesn't move your organization forward and it doesn't really help your customers. And I'm pretty sure if most people are anything like me, they feel pretty bad already and they probably don't need your assistance. This ultimately is just going to discourage transparency. It's going to make it harder to find out what went wrong. People will find ways to downplay what they did to preserve their self-image because they don't feel safe.
The other problem is if we excuse things as human error, we're not gonna have a conversation about it. If I just say that the dev on the team did this and they shouldn't do it again, there's really no reason to do anything else. Why would I improve the process? They're the problem. The reality is we work within the process, and the process is always part of the problem. If you do this, you'll have a broken system still. So a quick example of why we don't blame. If you come to GitHub and you see an issue that looks roughly like this. You're probably a little freaked out. You've done something, you've made some change in your code that's literally lighting computers on fire. It's kind of a bad day. But someone on your team's gonna sign it, they're gonna look into it, and then they're gonna just fix it.
That's cool This this might happen in a in a blameless situation, but it's almost certainly going to happen in a blameful one. Why did something catch on fire? What did we change? What happened? Oh, it's cool. They they knew what was wrong. They fixed it. It's great that we're not lighting computers on fire anymore, but we've done nothing to make sure we don't do it again. And if you're in a blameful environment, this is going to happen all the time. So how are you going to build this blameless culture? The first thing you need to do is create psychological safety for yourself and your coworkers. Because an environment where we feel that we're going to be blamed for anything that we do is an environment where we're not going to do anything.
There's a site that was released called Rework from Google that talks about team building. But this particular diagram is very much relevant for building blameless. uh environments. The most important thing before anything else is creating psychological safety. Until you do that, you can't depend on your team They can't depend on you. You're not going to have clarity. It's gonna be hard to build structure. It's also gonna be hard for anyone to have impact. Everything is going to derive from this. So until you do that, most everything you're doing is a waste The other thing you can do is find positive examples. I intentionally didn't bring up SRE, which is a discipline within technology, site reliability engineering. Because I think they already do this most all the time.
Their job is to make sure websites keep running, services keep running, and the rest of us can build new stuff. They help us build processes and help us avoid future incidents. So there are plenty of examples. This is just one. It's a book published by O'Reilly, written by Google Engineers about how site reliability engineering works. Uh it's a pretty easy read. I I thought it was very informative, but there are a lot of other examples in MMs and NTSB. I would encourage you to read more about this. The other thing, again, to remind yourself is outages will always happen. And no one on your team wanted them to happen. And if you find yourself ever wanting to blame someone, just recall that. So to close, I wanted to bring up a test for blamelessness of that incident we talked about.
You might start and just say that a script was executed and deleted the production database. Right? That's that seems fine. But interesting enough, uh, that's kind of passive. And someone had to run the script, it likely didn't run itself. And so an environment where you want to make the script run itself is probably kind of blameful. Maybe we'll get to the next place. An engineer on the team executed the script. This is pretty good. I mean at least you're saying someone did it. You are hiding uh information about who they are, but in some environments that might be good, right? If you don't think that broader than your team is going to feel this way. You might not want to put someone's name all the on sorts of docs that get distributed. But if you can manage to have a truly blameless culture, you can say that Chris executed a script that deleted the production database.
That incident we talked about was me, because I'm the only person on my team that's ever in the office that early in the morning. Uh and I did it by accident. We had a hundred branches that were stale and our GitHub pages branch that it was stale too because we hadn't released the documentation in a bit. And I deleted it. But I didn't have to worry about saying that because no one on my team is going to look at me and say, why do you work here? Okay. Someone is thinking this right now. This is a quote from an economist I like a lot that says, if you've never missed a flight, you're probably spending entirely too much time in an airport. If you're never failing, you're probably not trying very hard.
You know what you can do? You can break your own stuff. This is really fun. If you haven't done this, uh do this. I'm not going to go into a talk about creating chaos monkeys, but uh you should definitely look into it. The other thing is your SLA might be too low. If you're not breaking yourself, just throw some nines on there. You'll have some new problems. So uh thank you all. It was uh great to talk to you about this. Um I will take questions in the hallway uh or at lunch uh because lightning talks are gonna be starting next door. But
They can shift attention from assigning fault to understanding contributing factors and changing systems. Most importantly, they should treat prevention as more important than simply explaining what happened, using facilitated, non-adversarial discussions when necessary.
Discussed at 15:06Start with a concise incident summary, then document what triggered it, user impact, detection, resolution, and the detailed timeline. Follow that with what went well, what went poorly, what got lucky, and concrete action items to prevent a repeat.
Discussed at 16:42Action items can investigate the incident further, repair damage, detect future incidents, mitigate their impact, or prevent them altogether. They should be assigned real follow-up work rather than merely being discussed and forgotten.
Discussed at 22:08Blame discourages transparency: people hide or downplay mistakes, while the underlying process remains broken. A blameless approach makes it safer to share what happened so the team can improve the system and prevent recurrence.
Discussed at 24:29Create psychological safety first, so people can admit mistakes without fearing punishment; everything else, including trust, clarity, structure, and impact, depends on it. Positive examples from disciplines such as site reliability engineering, healthcare, and transportation safety can also help teams adopt better practices.
Discussed at 26:56Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026