EuroPython 2026 with Mia Bajić
Published May 6, 2026
This video features Mia Bajić at DjangoCon Europe 2025 in Dublin, Ireland.
Keynote: The Most Bizarre Software Bugs in History by Mia Bajić
https://pretalx.evolutio.pt/djangocon-europe-2025/talk/W3BQVT/
Mia Bajić uses bizarre software failures to show that disasters usually arise from interacting weaknesses rather than one isolated bug. She explains how the Boeing 737 MAX crashes involved flawed automation, a single faulty sensor, hidden system behavior, inadequate training, and delegated certification, then connects that pattern to Google’s Safe Browsing outage, NASA’s Mars Climate Orbiter unit mismatch, null and input-handling errors, integer overflows, spreadsheet mistakes, and a destructive Linux install script. The practical lessons are to eliminate single points of failure, make software behavior predictable and visible to users, define interfaces and units clearly, test integrations and configurations, handle real-world edge cases, and treat spreadsheets as production code—while recognizing that not every bug is harmful, since some become celebrated features or jokes.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: We've been all told we should test our software, but what happens when we don't? Sometimes it leads to strange, unexplainable events. Is testing war always the right solution? What do bugs reveal about software? And how can you use those lessons to build more resilient systems? Let's talk about the most bizarre software bugs in history It's 2018, the airport in Jakarta wakes up, the sun rises over the horizon, planes are heading to the runway, everything is as usual. Inside one of these planes, two pilots sit in the cockpit. Nothing usual. They have done this ton of times Together they've got over 5,000 hours of flying under their belts.
Speaker 1: They go through the regular plea flight checks, everything looks fine, instruments are good, weather's great, no issues at all. The plane heads out to the runway. Engines start roaring, it picks up speed, and they're in the air. Everything is smooth. A few minutes later, the pilot reached for the controls. Something feels off. The nose of the plane suddenly dips. They didn't do that. The captain pulls back on the controls and the nose lifts. For a moment everything seems fine. Maybe it was just a turbulence. Nope. It dips again. A warning flushes, airspeed disagree.
Speaker 1: What does that even mean? They have no idea what's happening there. They check the dashboards, they go through the emergency checklist, but neither of them really know what's happening. Going on there and they're 5,000 feet in the air. They try to pull the nose again, but the plane is pushing back. It feels like as if there was something something else or someone else flying the plane The plane races towards the ground at full speed and then at 11 pm 32 minutes Lion Air Flight 610 disappears from radar. What were the pilots actually fighting? Was it the turbulence? The engines? Or maybe just a few lines of code they didn't know about?
Speaker 1: To find out, let's go back to year 2010. As most of you already know, the biggest aircraft monocrafters are Airbus, which is based in Europe, and Boeing, based in the US. And they've been competing for decades. In 2010, Airbus made a big move. They announced the A320 NEO. You can see it on the right side. It was an updated version of their best-selling plane, A320, which is on the left side. And as you can see here, it has the same design, but it has just big, uh bigger and more efficient engines. What that means is that the fuel costs were lower and the operating costs were lower, so airlines loved it. Within months Airbus received thousands of orders.
Speaker 1: And Boeing had a problem. Their best sailing plane, the 737, was less efficient and more expensive to fly compared to Airbus. So they had two choices. One, they could build a brand new plane, but that could take for years because it's very hard to design and build a plane from scratch Or two, they could just modify the existing plane and make it more competitive. So because they were running the market race, they chose the second option So in 2011 Boeing announced the 737 Max. The idea was very simple. Just put bigger engines on it, just like Airbus did. But there was one problem with that If you compare these two planes, you can notice something, and that's that the 737 sits lower to the ground, so that means that you cannot just put bigger engines like Airbus
Speaker 1: did. and because they would just not fit under the wig wing. So how can you fix that? So they came up with an idea to move the air engines forward so they could fit However, what happens when you move the engines forward is that you shift plane's balance. Why? Because engines are very heavy. They're one of the heaviest parts of the plane. So if you move the heaviest part forward, the aircraft's center of gravity shifts. making it more likely to tilt. And you can see here on the first on the picture above, you can see where is the center of gravity, and on the second one you can see that the engine is bigger and it's more towards the left side, so the center of gravity is also shifted. Now, this necessarily isn't a problem. If you're a pilot, you know how to fly a plane that has a tendency to tilt.
Speaker 1: This is not a problem at all. But there is one catch. If you want to introduce a new plane which behaves differently, you have to train your staff. And training takes months, sometimes even years. And Boeing didn't like it because they were losing the market race. So they came up with a very interesting and a very unique idea. So instead of tr training pilots, why not add a piece of software That makes the plane behave exactly like the previous version. So they invented the maneuvering characteristics augmentation system. The way this work is that if a plane's nose tilts up The system corrects it by pushing it down. But there is one catch.
Speaker 1: Pilots didn't know about it So if you wonder how it is, imagine you're driving a car, you're going straight, and suddenly your car just drifts left and right, you end up in the opposite lane, and you have no idea what's happening. Uh and that's basically uh what they did. If you want to see, there is a simulation on YouTube where they put an experienced pilot go through the same conditions and he described it as quote unquote, it's a horrible situation, it's unimaginable Now let's take one more step back. How do you actually determine the position of a plane's nose? And this is a very important question because uh it's related to the software. In aviation, the industry standard is that all sensors measuring critical data, like speed or position, must be duplicated.
Speaker 1: And if both sensors give the same reading, like you can see here, that it is considered correct and it's displayed to the pilot However, if two sensors give different results, the pilot is alerted and follows a specific checklist of action to resolve the issue. And for each issue they have a specific checklist, they just go through it, so it's a standard process. Now, second problem is that this software relied on one sensor. The sensor called angle of attack was used to determine the plane's nose position and it was just on one side. Why? Well, we don't know. Um it's not in any official reports, definitely not an industry standard. So the best you can do is guess that it was some sort of um
Speaker 1: Management decision probably to reduce costs or complexity. So let's put all of these pieces together. What happened with Lion Air Flight 610? Boeing added piece of software to automatically adjust the plane's position if it tilted up too much. That software relied on a single sensor. That one sensor failed because it was miscalibrated, so it was sending incorrect data to the system. The software got activated unnecessarily and forced the plane's nose down, and the pilots didn't know that the software existed, so they didn't know how to disable it. Now some of you might be wondering, but how is this possible? Like
Speaker 1: has this plane ever been tested? So in March twenty seventeen, the Boeing seventy seven Max received certification from the Federal Aviation Administration, and the official statement said To earn certification for the 737 Max 8, Boeing undertook a comprehensive test program that began just over a year ago with four airplanes plus ground and laboratory testing. Following a rigorous certification process, the FAA granted Boeing an amended type certificate for the 737 Max 8, verifying the design complies with required aviation regulations and is safe and reliable. But after the crashes, uh investigators found something alarming. The plane was tested, but it turns out that the FAA
Speaker 1: delegated large parts of the certification process to Boeing itself. So that means That like Boeing was reviewing its own work. Um so critical safety systems were never actually evaluated uh by an invest by a third party um authority. It was Boeing engineers who signed it off as uh safe. After the crashes, the 737 Max was ground worldwide and in response they identified eight critical safety issues that had to be fixed before the plane could um fly again and Boeing was also charged With fraud. So Boeing made several changes to fix the problem that led to the crashes. Two sensors instead of one to mitigate a single point of failure.
Speaker 1: The software doesn't uh anymore activate repeatedly. Before this it worked in a way that it would repeatedly activate. So if you want to lift up the plane, you were not able to because the software was just uh activating every five seconds. Uh the movements are now limited, uh it just activates once and uh the range of the movement is limited. Um there pil pilots know that it it exists, so there is better pilot training, um there are new emergency procedures in case something goes wrong. Uh there is a warning that's always turned on. In the past uh there was a warning only if you pay for it. Um and and uh there are new sensor calibration checks to make sure that if sensors uh don't work That pilots
Speaker 1: know about it. And as you can see, it's never just one thing that causes failure in complex systems. And in risk management, there's a model called the SwissGee's model. So imagine you have slices of Swiss cheese with holes in them. One hole in itself doesn't matter because there is no hole on the next layer. But sometimes you get unlucky and all the holes line up and that's when the filler happens. So in this case uh the software had a flaw. The pilots were untrained on it, so there was some human error. And also the plane had only one sensor and that one sensor failed. And all of these by themselves might have not caused a disaster, but all of them uh together, that's what basically Basically made it deadly.
Speaker 1: So what does it mean to us? Some of you might be now thinking, well, I'm a Django developer, I don't really, you know, program software in planes, or I don't work on live critical systems where bugs kill people. But this might The model applies to all complex systems and you can find it in all major organizations. When there's some big failure, when something happens, it's never just one single thing, but it's always like a chain. reaction that multiple things happen at once that led uh to the disaster. And so why is it so hard to catch these um kind of failures like Some of you might say, but can you just not test all possible outputs? Well, it all comes down to combinatorics. So if you have one input, you have one output.
Speaker 1: If you have five inputs, you have 120 possible outputs. If you have ten inputs, you have over three million possible outputs. And this is all assuming that your system is deterministic. So think about modern software. How many lines of codes are there? How many people work on it? How much of infrastructure is involved there? From writing a piece of code locally to running it on the server, there are a thousand things that can go wrong. So what can we learn from a case like this one? First, single points of failure are dangerous. In this case, it was obvious they were just using one sensor and that one sensor failed, sending incorrect data. Failures are often a chain reaction. As you have seen, there are multiple things that happen at once.
Speaker 1: There were lots of flights with this plane, it was tested, but it was this specific case where multiple things happened Software should never do something users aren't aware of. In this case, PIOS didn't know that there was a software flying the plane they were supposed to be flying. But I think this is applicable for any kind of software. If you write software, don't surprise your users. They should be always aware what it's going and it should be very predictable for them. We've just talked about how failures, when combined, can lead to a disaster, but sometimes a single mistake is enough to cause worldwide disruption. Imagine this You're a developer at Google, it's a normal day, and you push a simple update to Google Safe Browsing, the system that warns users about dangerous websites.
Speaker 1: Just a small change, nothing unusual. And then you realize that every single website on Google is marked as dangerous, including Google itself. That's exactly what happened in 2000 oh in 2009 when a Google engineer made one tiny typo. For 40 minutes every Google search result said this site, my harm or Computer and even Gmail was affected. Uh the best part of it is there was just one service that was not affected, and that's the ads. If you're wondering why, I have no idea. Always work, they're always there to spam you. So what went wrong?
Speaker 1: One of the engineers was updating the malware registry, which is a list of dangerous websites, and instead of entering a specific URL, they accidentally added just one character, a forward slash. Which means everything. So what we can learn from this is that typos happen and sometimes testing more is the answer. But not every bug needs fixing. Some are so iconic that they never get fixed at all. Who here knows the game's civilization? Oh, lots of people. So it's a strategy game where you build a growth civilization and you're competing against historical leaders like Napoleon, Cleopatra, or Genghis
Speaker 1: Khan. And then there's Gandhi. He's a leader. Known for peace and diplomacy. And at first everything seemed normal, Gandhi preferred negotiations over war, just as you would expect. But as players progressed through the game and civilizations became more advanced, they started noticeing. Something strange. The moment Gandhi unlocked atomic technology, he started dropping nuclear bombs off everyone. So why? Well, the urban legend says that the Developers gave Gandhi the lowest aggression rating possible, a one. And later in the game, when civilizations became more peaceful, the game automatically lowered aggression scores by two. So what happened binary
Speaker 1: is one is zero zero zero zero zero zero one. Subtracting two made it roll to zero and then to one one one one one one one one which is two hundred and fifty f the highest aggression level possible so instead of a peaceful diplomat Gandhi became the most war-hungry leader in the game now Now there is one catch actually. Although it's often cited as a bug, the game designers and developers said that it's not a bug. It started as a joke on Reddit and it became viral. People loved it so much that this Designers decided to introduce a bug because it was getting so popular. So in the latter versions there was Gandhi dropping bombs and everyone turning a myth into a joke.
Speaker 1: So what can we learn from this? Well, not all bugs are tragic, and sometimes they become leg legendary Easter eggs that can even boost popularity. And sometimes a bug doesn't become a feature. Sometimes it has become a very expensive mistake Like the one that cost NASA three hundred twenty-seven million of dollars. It's September nineteen ninety-nine We are at NASA 's Jazz Jet Propulsion Lab watching the Mars climate orbiter finally reach its destination. Mars climate orbiter was a robotic probe, as you can see on this image. It's about the size of a small car and its job was pretty straightforward. Cruise to Mars and orbit the planet to study Martian weather and climate.
Speaker 1: Think of it as some sort of uh weather satellite for Mars. There were no astronauts no astronauts on board, no staff, it was just lots of um sensors and lots of electronics. The team had been working it for years and it cost hundreds of millions of dollars to bid on and launch. So back then it was pretty much of a big deal for NASA. After 10 months of traveling The arbiter is finally reaching its destination. Is it Pomin Day and everyone's been waiting for and people are watching it on TV So what was supposed to happen is that the orbiter was supposed to enter the orbit and to come from behind the planet and radio back and say, hey, I made it. But instead, there was silence Minutes pass, hours
Speaker 1: nothing. The Mars climate orbiter had it essentially vanished. So what happened? At first, no one knew. It was supposed to be 110 kilometers from Mars But then they realized that instead it had dropped to fifty-seven kilometers, so low that it was basically grazing the mountain tops. And at that altitude the Martian atmosphere tore it apart. It probably uh burned up or broke into pieces. Completely destroyed. But why? What could have pushed it so far, of course The spacecraft was built by an external contractor on the left side. But it was operated by NASA's navigation team, which was a different team.
Speaker 1: And there was a piece of software that calculated the force of thruster firings, and it was integrated. with a navigation software that was reading the data. So one software, one team was generating data, the second one was reading it. And when they checked the numbers, the numbers were right But the meaning of those numbers was off by a factor of four point forty five, and that's when they realized where the problem was. The first software output those values in pound for seconds, which is imperial measure of impulse, but the navigation software reading the data assumed it was in Newton seconds, which is metric. So over the course of the mission, every time the orbit Small thrusters fire to adjust its course or orientation, the recorded impulse was misinterpreted.
Speaker 1: So you might be wondering, um, how does such a basic error slip through for a mission this important and at NASA? Um it's a great question. Uh it was mostly obviously a communication failure because NASA's navigation team assumed that everything was in metric. Um and there was a report where it was written that everything is metric and it was signed by both sides, but probably not. No one read it or I don't know. Um and the worst part of it is that um s the data had shown consistencies be weeks before the failure, and some people tried to raise it up, but uh these weren't fully investigated. It was probably probably management decision uh not to investigate it. So what can we learn from this? I think this is actually a great example because even though it's NASA
Speaker 1: and spacecrafts, uh so many things are applicable to any kind of software. So, first thing is establish clear interfaces. If you have two systems and two systems interact, make sure they agree on formats, units, and overall assumptions. And here we have two systems interacting together and operating under different Assumptions and that's uh what basically happened. Uh it here it was just one bad assumption that broke the whole system and a huge mission that cost so much of money and time. So sometimes even small mistakes uh like this one can. And have huge consequences. Integration tests are useful. If NASA had simulated the orbiter's navigation with the actual data, they might have caught a discrepancy before launch, but they never had any integration tests.
Speaker 1: and uh learning improved processes. After this disaster, NASA standardized units across the organization, enforced stricter checks and uh improved interteam communication protocols. In the mid-90s, a new employee at Sun Microsystems in California kept mysteriously disappearing from the database. People started investigating only to discover that the issue wasn't the system failure. It was his name. His name was still Steve Null. It turns out that some systems back then didn't still handle the string uh null properly. I tried to reproduce it in pro Postgres locally, um
Speaker 1: created a database, instead inserted the value now as a string, and if I tried to retrieve it, as you see it works fine. There are bugs like this in various systems, mostly older ones. Uh for example, I found an open issue in Apache Flex. Um it's a it's an old software that isn't really used anymore, but uh some time ago um it was uh it was an issue in many Systems. And we've all heard the urban legend, I have no idea if this is true, about someone who changed their license plate to drop database and attempted a SQL injection attack. Speaking of SQL injection attacks, I cannot not mention that popular comic um Hi this is your son's school, we are having some computer trouble.
Speaker 1: Oh dear, did he break something? In a way, did you really name your son Robert Job Table Students? Oh yes, little Bobby Tables we call him. Well, we've lost this year's students records. I hope you're happy. And I hope you've learned to sanitization. your database inputs. Some automated systems will delete all records that start with TAS or ABCDE just because they assume it's all TAS data, even though they're not. If you're now thinking, wait, you said ABCDE, like who would give their kids such a name? Well, actually, I have some news for you. Between 1990 and 2020, over 300 babies in the US.
Speaker 1: were named ABCDE. I can only imagine how well this goes with modern software. So what can we learn from it? Handle edge c uh edge case in user inputs because real world data can surprise you For a long time we thought that the most expensive book ever sold was the Codex Licester, written by Leonardo da Vinci in the 16th century. It was later bought by Bill Gates for thirty point eight million in nineteen ninety four. A book written by Leonardo da Vinci, uh written by the thoughts of one of history's greatest uh minds, lots of handwritten pages of Scientific observations and ideas that were centuries ahead of their time.
Speaker 1: So, what kind of book would have to be to be so expensive? Some rare manuscripts, some lost work of literature, something of massive historical importance. Well, in 2011 a book appeared on Amazon with a price tag of twenty-three point seven million of dollars. And no, it wasn't some ancient text or some rare manuscript, it was a book about the genetic development of flies. And funny enough, just like the bug on the cover, the price was a bug too. So here's what happened. On Amazon you have lots of sellers who can set their own prices. And uh they have automatic pricing groups that adjust their prices based on competitors
Speaker 1: So one seller set their rule to something like always be 0. 07 % cheaper than the next lowest price, and the next one had something like always be 27 % more expensive than the lowest option. Normally this works fine if you have lots of sellers, somehow the prices will balance themselves. But here the catch was that there were only two sellers and their pricing algorithms got stuck in a loop So every time the system updated, one seller's price went up and the other one followed even higher. And again, again, until the book was listed for twenty-three point seven million of dollars. As soon as people noticed the fist the The system was fixed and the price dropped back to normal.
Speaker 1: So what can we learn from it? Well, be careful about your assumptions Are field can sometimes seem mysterious to outsiders. When WhatsApp said the maximum number of people in a group chat to 256, the end independent reported it with a comment. It's not clear why WhatsApp settled on such an oddly specific number. That comment quickly disappeared and it was replaced with a footnote. A previous version of this article said It was not clear why WhatsApp settled on the oddly specific number. A number of readers have since noted that two hundred and fifty six is one of the most important numbers in computing. Since it refers to the number of variations that can be represented by eight switches that have two positions.
Speaker 1: eight bits or a byte. This has now been changed thanks for the tweets DB. Speaking of 256, did you know that trains in Switzerland are not allowed to have 256 axles? To keep track of trains on the Swiss rail network, detectors are placed along the rails. These simple detectors activate when a wheel passes over them and count the number of wheels in order to provide basic information about The plane. Unfortunately, these detectors store the axle count in an 8-bit binary number. So when the count reaches 11111111, adding it one rolls over to 8800000, which means that the that the The count is reset back to zero and the train becomes invisible.
Speaker 1: So it's a phantom train moving through the network undecat un undetected. And if you check Swiss railway regulations, you will find section three point seven. Which states to avoid falsely signaling a section of track is clear by resetting the exil count to zero and thus to avoid collisions, the total number of axils in a train must not equal 256 Some two hundred fifty-six related errors are harm harmless. Some are funny. And some are deadly. The third 25 was a medical radiation machine used to treat cancer. It had two modes A low power beam for regular treatment, which you can see on the left side, and a high power
Speaker 1: beam that had to pass through a metal target to turn into x-rays. The high power beam was super dangerous if fired directly at the patient. So the machine had a safety check and relied on software to make to uh make sure the metal target was in place before use. Funny though, uh the variable tracking the state of the check was called class 3. And if class 3 was zero, that meant that the check was run, the target is in place, and the beam can be fired. But there was one problem. As you can guess, the variable was stored as an 8-bit number, so meaning um every 256 cycles it rolled over back to zero even if the safety check never ran And one day that's exactly what happened.
Speaker 1: In 1987, in the Yakima Memorial Hospital in Washington, a patient was supposed to receive 86 rats of radiation. But at the exact moment the operator hit the button, the variable class 3 had rollover from 256 to 0. The machine skipped a safety check and fired the high power beam straight at them. So instead of 86 reps, they got a deadly dose. Of 8,000 rats. A simple 256 rollover bug turned a life-saving machine into a deadly one. Sometimes it's hard to understand what is a bug and what is a feature. Take Excel, for example. Everybody uses it, even biologists. In 2016, researchers in Bellburn
Speaker 1: analyzed 18 genome research journals and found that 19. 6% of gene studies contained errors. How is that possible? Well, did you know that there are genes called March 5 or SEP -15? Well, Axel didn't, so it autocorrected them to a date. So if one in five gene research papers had axial-induced errors, how about it is in the business world? For that, we turn to the European Spreadsheet Risks Interest Group. It's an organization that studies spreadsheet mistakes. And they estimate that around 90% of business spreadsheets contain errors. So how did they get that number?
Speaker 1: It sounds very specific. In two thousand one, during the Enroll scandal, the US government released half a million from the company with spreadsheet attached. So researchers studied these emails as real world snapshot of how spreadsheets are actually used and they came came up with some very interesting numbers. The average spreadsheet size was around 114 kilobytes. The average spreadsheet had around 6,000 non-empty cells. One spreadsheet had 174 worksheets. 6. 6 thousand spreadsheets contained no formulas at all, and 24 % of spreadsheets with formulas had at least one error. And sometimes those errors cost
Speaker 1: lots of money. JP Morgan used spreadsheets to calculate risk. Specifically, there is something called value at risk. It's supposed to answer the question, what's the most money we could lose in a 95 % certainty? It's based on historical price movements, volatility, and current portfolio positions. The way it works is that if the value at risk is low, traders are allowed to make bigger risks, but if it's high, then they should play it safe. And one specific value risk calculation was being done in a series of Excel spreadsheets with numbers manually copied and pasted between them. This was supposed to be a prototype, but what happened is that someone put it into production. So the risk calculation was wrong.
Speaker 1: And the valuation control team, which is a separate team, and their job was to check if trader like marking their portfolios correctly. They were also relying on Excel. And those spreadsheets were also wrong. And at one point it got so wrong that one that one employee started their own unofficial spreadsheet just to track actual profits and losses. And by the time They caught a mistake, it was too late. JP Morgan later published a report detailing what went wrong. After subtracting the old rate from the new rate, the spreadsheet divided their sum instead of their average. This error muted volatility by a factor of two, lowered the value of of at risk. What does it mean is that JP Morgan lost six billion of dollars because someone added two numbers together it instead of averaging them.
Speaker 1: So the biggest thing we can learn from it is that spreadsheets are code as well. So you should treat them like production software. We've seen how software bugs can cost billions. They can destroy spacecraft and even turn medical devices deadly. But sometimes a bug doesn't wipe out a company or cause a huge disaster. Sometimes it wipes out your entire operating system. Bumblebee was a project designed to enable Nvidia dual GPU support on Windows laptops, allowing users to offload graphics rendering to the discrete GPU. Everything was working fine until one update in 2011. A user created a GitHub issue with a rather urgent message.
Speaker 1: Install script does rmrf user for Ubuntu. An extra space at line 351 causes the install script to do an RMRF on the user directory for people installing it Ubuntu. Totally uncool, dude. The script deletes everything under user. I just had to reinstall Linux on my PC to recover. Removing the extra space will fix this. Possibly do it quickly. And the author quickly pushed a fix with a comment message that really summed up the situation. Giant bug causing user to be deleted. So sorry. And of course this being open source the community did not hold back in the comment session. How can you complain about bugs, Mr
Speaker 1: Anderson, when you have no operating system? Bleeding edge really bleeding for someone now. No more lack of disk space now I didn't like that folder anyway. I 'm here by thank you for your attention. I put together a list of resources if you would like to go through it and I also uh uh put here two favorite books about uh software flaws. One is Humble Pi by Matt Parker. He's an Australian British
Speaker 1: stand-up comedian and uh programmer and uh and mathematician and he wrote a great book about lots of very bizarre stories. And the second one is the design of everyday things. um which is about design in general and how people use things uh in a way that they were designed, I totally recommend so if you're reading in if you're into reading and um if you have any other suggestions for books, I'm always happy to hear them. And that's Now it's time for your questions. You
Speaker 2: told lots of stories, but um none of them was about yourself or your own experiences. So what's the closest you've been to uh uh A bug or an error that could have appeared in a presentation like this?
Speaker 1: I knew that this would be the first question. Well at each job I signed an NDA contract, so I'm not sure if if it's a good idea to reply to this question. Yeah, I don't have any similar I I don't have any such bizarre stories definitely, but I don't have any stories that I can share publicly. But if anyone else has
Speaker 3: Are you perhaps interested in in more uh bugs because I once uh mm made a path traversal bug that uh deleted my colleague's home folder and stuff like that? So uh are are you interested in more stories? Uh should I come to you later or
Speaker 1: oh I am we have ten minutes so
Speaker 3: Yeah, there was also one other bug in in uh in an Excel sheet that I found interesting that was when when uh I uh uh imported a CSV. with a uh comma uh and three digits after it and uh somehow excel decided that it was a thousand uh point so um it was not uh the Comma point that I expect expected it to be and uh oh yeah about the um about the uh units getting wrong. Um I'm really interested if there's any uh official uh Django um option to put a unit next to a float field because I use it in my production
Speaker 3: uh that I I extend the float field to add this uh This um format uh wait, the uh the unit uh value to it.
Speaker 4: I think um a lot of the the bugs you showed were quite entertaining or interesting, but one thing in common is that there was human error involved in all of them at some point. Um with the rise of AI generated code perhaps making its way into production more often. Do you have any thoughts on what your talk on interesting bugs might include in ten years time.
Speaker 1: Yeah, I think that's a really good thought. Ten minutes it could be the most bizarre AI-induced software bugs
Speaker 5: Hi, thank you. Great talk. Um do you think I've uh all of these problems could have been solved with more testing or Is it one of those problems where everyone feels like they've tested it and don't know until these things happen and it's only in the aftermath that they make the real learnings? Um to have an opinion?
Speaker 1: Yeah, so I think it depends. I think that definitely some problems could have been solved with more testing. So for example the the example from NASA where we had two softwares that were using different units, they had no integration test. So obviously if they had some integration tests they would have known that they are using different Unix. So that's an obvious one. Or the ones from Google, for example, about the malware registry, they also had no tests. So yeah, but I also think that um like we are introducing new technologies and systems are getting more complex, and I think that you can never test all possible um combinations. Uh look at, for example, some very complex systems like aircrafts, how many buttons and different things are there? You test some of course you test it, uh but then you have so many variables you have human errors, you have mechanical errors, weather that's also a big thing
Speaker 1: Um so I and especially when we are like introducing new things, you you don't you you have a lot of the unknown unknowns, you don't know about things. So I think that that's part of the progress that's Like you introduce new things but sometimes just errors happen and no matter how much you test, you cannot test out everything And that's actually it's a funny story. It reminds me of a story about planes. Why do planes have no why planes have square windows? Well uh well because in the past they they they started experimenting with making square windows because customers wanted it because it looks nice or whatever, but then they realized that the problem is that the planes at the beginning were made of wood and if it's made of wood it's not a problem
Speaker 1: because you know wood is very resistant and everything but then if you use metal uh And you have a plane with scare windows, you have lots of pressure in the corners. So what happened is that those planes, those windows would be in the air and after many hours of flying they would just break apart And that's because it's a new thing. Metal was not used before that, before that they used uh wood. So I think as we are discovering new things there'll be always like uh cases like those
Speaker 6: So uh the most of the examples you've shown were basically uh software bugs. But the uh Google example was it's uh seems like it was actually a configuration bug. And and those are basically different. Can you comment on on testing configurations and how to make those maybe more resilient?
Speaker 1: So my question to you is what is the difference between software and configuration? What do you call software and what do you call configuration? Because for me it's like soft configuration like software is an umbrella for for configuration for uh lots of programming different things in uh or at least in this case but maybe try to define it more
Speaker 6: well Configuration is something you do without changing the actual binaries or or the the code that runs. Um or or even maybe similar error to that Google or uh there was a BPF Like network configuration change in in Facebook a few years ago that brought all those systems down for like a couple of days Uh so so that's what I think of as as configurations. Settings that are not affecting the question.
Speaker 1: Yeah, so the question is what is the difference between testing the th two or what was can you repeat the question please?
Speaker 6: Can you comment on on how to make configuration changes um more testable and resilient as opposed to code?
Speaker 1: Yeah, I think that's a very broad question, so it really depends what kind of configuration change are we talking about. Uh in this case I know that they didn't have any test for this particular thing and afterwards they added tests but I think this is a very broad question to now reply in one sentence but if you want we can uh talk about it later um if you have any specific examples
Speaker 6: Thank you very much.
Speaker 1: Thank you.
Speaker 7: Thank you so much for keynote. I thought it was amazing. A couple of times you made cheeky comments about management making cost decisions. But there is a cost trade-off Right, as you scale up the testing it become it's not they're not always just being
Speaker 1: tight fisted, are they? Well, everything is a trade-off, so But this is a hard one because uh there are lots of decisions to be made and when everything works no one sees that like everything works and there was some decision made but sometimes people make bad decisions and they make trade-offs and Then disasters happen. Thank you. Thank you.
A miscalibrated angle-of-attack sensor sent incorrect data to the MCAS flight-control software, which repeatedly forced the nose down. The pilots were not told about the system and did not know how to disable it, while the plane’s certification process had not independently evaluated the safety system.
Discussed at 7:21Major failures usually result from several weaknesses lining up rather than from one isolated bug. In the 737 MAX case, flawed software, a single failed sensor, and insufficient pilot training combined to create the disaster.
Discussed at 10:26The number of possible outcomes grows combinatorially as inputs and system complexity increase, and real systems also involve human, mechanical, environmental, and unknown factors. Some failures can be caught with targeted or integration testing, but no amount of testing can cover every possible combination.
Discussed at 11:57An engineer accidentally added a forward slash to Google’s malware registry instead of a specific URL. The slash matched every website, so Google’s systems warned users that every site—including Google and Gmail—was dangerous.
Discussed at 14:22The story is that Gandhi’s very low aggression value underflowed when the game reduced it, wrapping around to the maximum value and making him launch nuclear attacks. The developers later embraced the popular bug—or legend—as a deliberate joke and Easter egg.
Discussed at 15:08One system reported thruster impulse in pound-force seconds while the navigation software interpreted it as Newton-seconds. That unit mismatch gradually sent the spacecraft onto the wrong trajectory, causing it to enter the Martian atmosphere too low and be destroyed.
Discussed at 19:00Some older systems handled the string “null” incorrectly, treating Steve Null’s name as a missing database value rather than ordinary text. As a result, records associated with him could appear to vanish.
Discussed at 21:23Railway axle counters stored the count in an 8-bit number. At 256, the counter overflowed back to zero, making the train appear invisible to the signaling system and creating a collision risk.
Discussed at 26:52A safety-state variable was stored in an 8-bit number and rolled over to zero after 256 cycles. Zero indicated that the safety check had passed, so the machine could fire its high-power beam without the required metal target in place.
Discussed at 28:26JPMorgan’s risk spreadsheets divided a sum instead of averaging it, understating volatility and the value-at-risk calculation. The resulting mistake contributed to losses of about six billion dollars, illustrating why spreadsheets should be treated like production software.
Discussed at 32:17Yes, in some cases: integration tests would likely have exposed NASA’s unit mismatch, and tests could have caught Google’s malware-registry error. But increasingly complex systems contain too many combinations and unknowns to test exhaustively, so testing cannot eliminate every failure.
Discussed at 39:04Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025