Finding 2.0 with Marc Gibbons
Published December 6, 2024
This video features Marc Gibbons at DjangoCon US 2023 in Durham, North Carolina, USA.
Have you ever encountered a codebase where modifying code seemed impossible due to the constraints of the test suite? Are the tests that you write today inadvertently restricting future improvements? In this talk, we will uncover how placing empathy at the forefront of test suite development can empower rather than hinder future developers, including your future self.
This talk was presented at: https://2023.djangocon.us/talks/empathetic-testing-developing-with-compassion-and-humility/
LINKS:
Follow Marc Gibbons 👇
Follow DjangCon US 👇
https://fosstodon.org/@djangocon
https://twitter.com/djangocon
Follow DEFNA 👇
https://www.defna.org/
Video production by the presenter and DjangoCon US 2023 volunteers.
Marc Gibbons argues that tests should protect a system’s behavior and users’ outcomes, not the particular implementation used to produce them. Using a loan amortization example, he shows how direct imports, request-factory tests, private-method assertions, and mocks can create implementation bias: tests pass while the application is broken, then fail when harmless refactoring or dependency upgrades fix it. He recommends testing through real URLs and user-facing interfaces, using tools such as PyQuery, responses, and freezegun where appropriate, and treating a conceptual feature as the unit under test. The broader argument is that empathetic, humble tests make future change safer and kinder to both teammates and oneself.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Okay, so uh as I was saying, I live uh outside of Toronto in the other Durham, so it's been some interesting conversations uh about where I was going this week I've been writing Django since year 2012. Before that, I used to play professionally in symphony orchestras. And now I use Django mostly to help pay for my daughter's competitive dance bills. So we'll get started. Empathetic testing. I'm going to share a story with you today about something that happened to me back in 2016-2017 I was working at a fintech startup and the company was trying to fit find its like product to market fit and it started out as this like loan matchm uh this matchmaker right for um investors and alternative lenders And it started to gravitate and shift towards more of like an analytics
platform. So we as web developers were tasked with implementing these financial projections and data models built by the data science team. And in doing so, we came to rely heavily on libraries like NumPy and Pandas. And so as I became more familiar with these libraries, I realized that a lot of the code that we'd written ourselves Could be replaced by some of the financial functions that were provided in NumPy specifically. And I had a specific area of the code base in mind. It was our loan amortization schedule generator Now, a loan amortization schedule looks like this, right? It's a list of payments over the course of a term of like your mortgage or a loan. And each payment is broken down into two parts, an interest portion and then the amount that's applied to the principal outstanding
And so NumPy offered these two functions, PPMT, IPMT, and I thought, great, I can delete a bunch of code, I'll use these, and you know, less code is better code, right? So, what did I do next? I ran our test suite, which I was very proud of at the time, had like, you know, was an iota away from 100% coverage. And what happened next would be like this inflection point I was telling you about in the opening. The test blew up. Why? I mean the math was exactly the same, the outcomes were the same, and the errors I were getting weren't exactly helpful. I was getting an import error, an attribute error, a name error. And the most egregious of all was this
exception not raised, like this stop iteration wasn't raised. And this is real code, by the way. And like I was super confused, like what does this have to do with my schedule? So I got frustrated, right? I'm coming along, I'm trying to make this change, this improvement, right? And I'm being held back. My tension, right, and my timer being hijacked. Meanwhile, I had other things to do that day, right? I had meetings, I had deadlines, I had features to build, I had pull requests to review. But I really wanted to get this done. And so that frustration and that uh that this you know that dissatisfaction boiled over into anger. Show of hands like if you've like experienced this before. Like, yeah, yeah.
Like maybe you're even like experiencing it a bit right now, thinking about it. Uh-huh. So, what do we do in these situations, right? I thought I had the perfect solution. I thought at the time, what if I could displace and project my anger, my frustration, my dissatisfaction? onto the author of these terrible tests. I had to know who did this to me, right? And luckily we have an EPI for that. Shout it out if you know what it is. Yes, thank you. Give blame. The solution to all our problems. So out of self-righteousness and contempt, I ran give blame. And do you think this solved my problem? Like, do you think it made my test any better? No. And I didn't even feel better about it.
Anyway, hold on to that feeling and we'll we'll come back to it later. So I've prepared this sample project today, which kind of recreates the scenario I found myself in back in the day. And the first example I'm going to show you is this REST API, which exposes our amortization schedule calculator. It's at the path slash calculator. It takes three inputs, our principal amount, the interest rate, and the number of periods, and it gives us our schedule. I have a test suite and it's got 100% code coverage, right? Yay me. But you'll notice that we're on version 3. 6 of Python because the year, as you remember, is 2016. So the first order of business was let's upgrade to the latest version of Python, which until two weeks ago was 311, is
when I made these slides. And we'll bump all the ver various dependencies and we'll run the tests again. And wow, everything passes. Good job, right? The best upgrade ever. So let's go to the browser and refresh the calculator and confirm that everything works. And lo and behold, the yellow screen of death, right? Why? My test pass. But at least this time I've got like a pretty clear error message, right? Like like NumPy has done a really good job of telling us like what the error is and more importantly how to fix it. And in this case, IPMT has been moved out of NumPy and uh put into its own library, NumPy Financial. So I went ahead, right, I installed the library, I updated the references, and I refreshed the page, and
we're back in business. Now what about our test? Do you think they pass? They blew up. So here we have like the exact opposite of the desired outcome, right? I have a test suite that passes completely defending this broken application, and then the instant that we fix the application, our tests are broken. Why is that? And I thought a lot about this in preparing this talk, and I this is what I came up with. I called it the implementation bias. And what I mean by that is that our tests aren't actually like testing how things work. They're testing the way the thing was built, right? And that means that we can't make any changes to it, right? It's completely immutable So here's an example of that implementation bias.
It's a test for that view that I showed you, the calculator. And there's a number of things wrong with this, and you might see it already. But the first thing it uses the request factory And I'll show you why that's a problem. Well, the first problem is this is the wrong URL, yet this test passes, right? The next problem is I'm importing the view directly and then I'm passing this like wrong request into it. And what that does, it really forecloses any opportunity for us to refactor. Like say right now I'm using this function-based view, but I'd like to change that to a class. Right, this test would fall apart and the functionality might still be the same. So that's kind of a problem. And then finally, right, right, we've patched this function, get amortization schedule. What does that mean? I don't know. And we're we're doing some, you know, asserting that we've called the function with some arguments and
we don't really know what we're doing. It's not clear. So what if we wrote a better test? What if we bent back and put yourselves, ourselves, in the in the shoes of that frustrating situation when we come to make an improvement and things fall apart? What if we use a bit of empathy, right, to prevent that situation? So the title of the talk is, you know, empathetic testing, developing with compassion and humility. And what do I mean by that? Like empathy. Empathy is like feeling the feelings and like relating to that situation that we've all found ourselves in. The next one, compassion, is kind of rooted in empathy, but there's a call to action. It's like, I see you over there refactoring and suffering, and I'm gonna do something about it, right?
I'm gonna prevent that from happening. And then finally, humility. Humility in this context is about accepting, if not embracing the fact that things change, right? We're not building things to a state of completion and then moving on. Right, libraries change. IPMT moved out of NumPy, right? Um your implementation of today, right, might have been fine and suitable, but you might have slept on it and thought of a better way and the next day Right, you might want to change it. Or somebody comes on your team and says, hey, I could do this 10 times faster with 10 times less code. We want to create an opportunity for them to make those changes. And humility here basically means that we have to treat our production code as though it is discardable, right?
It's throwaway. So let's go back to our calculator and put ourselves in the shoes of the user. Right? What are we doing? We're sending some inputs and we're expecting some outputs. So why don't we just write our tests in that way? Right? So this is much better. We're using like a client and we're doing a real request to the real URL that's exposed and we're asserting like that the data return is the data we expect. And that's so much more like comprehensible, readable, and it works. Now you'll say, hey Mark, yeah, you only got 96%. Like that's that's no good, right? So let's take a look at the coverage report and you'll notice that you know we're not Tracking this like condition that returns early if we have a false C value for a number of periods, presumably this is like to prevent a zero division error.
But let's just write a test for that and capture that test case. We'll give it zero periods and we'll get our test and our coverage. This was so much better, wasn't it? The next example is in the Django admin. So a little trickier to test, right? I have a loan model, right? It has the three fields: principal, rate, and number of periods. with the values that we've come to uh to um familiarize ourselves with. And here I'm rendering out the AMWort schedule in a table Now here's a test that has like a high bias towards the implementation. And I'll start right out here is the loan admin, right? We're importing the class itself. And there's an issue with that, right? Like, do we know if that's been registered to the admin site, right? And if we were to pull back even further, is the admin
site like is admin even installed in Django, right? Like is it in installed apps? Is the URL exposed at like slash admin or bananas or something other? Something else? Right? So that's not great And then you know we're creating this loan with no values. We're testing what appears to be this like private method. And then again, here's our mock that's kind of cryptic and unhelpful. So let's do something better, right? Let's focus on um what happens as a user, right? So we'll create a super user and log them in. We'll create a loan, a real loan in the database with the values that we you know expect. And then here's our schedule, and we'll actually call like the URL that's exposed in the system
and assert that it's a 200. But now you're thinking, well, Mark, this is HTML and this is different, and this is what we have to deal with. But you know, we we can manage with this. And there are libraries that can help us out. And the one I'm going to show you today is called PyQuery. So what we can do is import um PyQuery. We'll load our response in there. So we have like this DOM object that we can use a query selector on. We can grab our table and then we've parsed out all the noise. And then we'll like massage that data a little bit. And then why not load it into a data frame? Because we're data scientists now. And then we've got our data in the format which we can perform this assertion against, right? And that works. The last example in this fake project I've created
is another API. It's a loan creation endpoint. So it does two things. You post the data, it creates an object and saves it in the database. And then upon save, it sends an HTTP request to this fake notification service. And it's going to send information about the loan. And the schedule is going to be one of those fields that's sent. So our test that we're starting out with isn't a much better place, as you can see. Right? We're using a real client. Uh the real URL, right? It's correct URL. And then we're getting like the real object that was created from the database using the ID that was returned in the response. And then here's our schedule that we expect to be sent, and it's sent some other pieces of data. But this could be better, right? You'll notice that we're using, yeah, there we go.
Uh you'll notice that we're using uh a patch on the request library, which is what makes the HTTP call. But this this isn't great because There's an implementation bias right here. You'll notice it says like loan. views. requests. Well, that's the location in which the request library was imported. So that like eliminates the possibility of moving things around. Like say your views. py got really big as often does, and you wanted to split it up into smaller modules, this test falls apart. The other issue is there's a production bug that's being hidden. I don't know if you can see it, but you will in a minute. So what if we could um and that like the spirit of the mock is obviously that we don't send like live network requests. over the internet when we run our tests.
So what if we could eliminate the requests patch and still not send data over the internet when we run our tests? And the answer is we can, right? And there's a number of libraries out there, and the one I'm going to reference today is called responses. I see a few heads nodding as in you've used it. And um And the way we use that is we we can decorate our test or use it as a context manager and we'll register the URL that we uh we expect to be called out. And then once we've executed the code that makes the request, we're able to inspect what was intercepted. So here I can assert that we've called this notifications. test. The request method was a post, and here's the data that I expect it to be sent to our service.
So now let's run the test. Okay. Here's the error, right? Date time is not JSON serializable. And so this error was being completely like obscured by the fact that we were using uh a mock, that we were patching the request library. We're now executing more of the stock and we're getting more information. So we'll figure out where the date time came from. And then we think like simple solution is to like convert it to a string. So ISO format puts it in exactly what it sounds like. And we might be inclined to do the same thing on our view as well as our test, but this is kind of a like a weak assertion. Like this is like saying, assert true, really. So where does this come from? Create a daytime
comes from our model. It's a daytime field with auto now add, which means that you know we create the instance and it gets timestamped with the current time. And that's hard to test, right? Because every time we run our tests, right, the time changes. And like if you were to capture your the time, like the time the test starts and the time your assertion is made is different, like microseconds apart. So you could mat you could like mock timezone. now or do some funny things or um you could find a way to like freeze the time and there's a library out there that helps us do this called FreezeGun and Forgives a violent name, but it comes with a function called freeze time that you can use a decorator or as a uh context manager. And so we'll freeze the time at today.
And then in our assertion, we can get rid of the reference and actually assert an absolute value, like the string that we expect to be sent in the format that we've come to expect. And what's nice about this is that like we don't care how that came to be. It could have come from the auto now add uh functionality from the field on the model. Or we could have like said time zone. now in the view. We don't care. And testing in this way has allowed for both those possibilities to exist. So this leads me to refactoring, kind of where we're going with this talk, the goal, right? To avoid that situation where we've refactored and things blow up, but the functionality is the same. This is the code for the amortization schedule calculator.
I left it purposely small so you can't read it because it doesn't matter. There are two things to take away from this slide. And they are. There's a lot of code here. It's split up in all these like little functions. And basically the destination for all this is this trash bin. And we'll replace it with something completely different. And again, it doesn't really matter what it is. What we want to demonstrate is that our test suite continues to function, right? We're able to edit production code, but not test code. And bonus here, we've deleted a bunch of code, so which I still think is better. And so we'll run the test or three tests that we've implemented with this sort of like empathetic mindset to the refactorer, and we've got our coverage on that calculator, and we're not testing it directly.
So let's go back to that series of tests that we um encountered at the beginning once we did our Python upgrade. We have even more failures now, you'll notice. And it's these kind of lousy attribute error. What does that mean? And basically it's telling us that these functions that we're testing can't be false, found. And really, this is just a big case of implementation bias. These tests didn't really accomplish anything. They didn't test our functionality. They just defended somebody's opinion at a certain moment in time and made our life difficult. So we've got a target for that. We're going to take that entire test module and we'll put it where it belongs. So you might be thinking, well, Mark, that's all great.
But you've kind of repeated yourself, right? Three different places that use the schedule. And you've only really shown us like one scenario, this $10,000 loan, which is paid annually, which loans don't typically work that way What about all these edge cases? Like we captured the zero period, but only in one instance. And so you might also think like that's the whole point of testing this unit, the schedule, right? That function. And hammering it with all the edge cases and not really concern ourselves about the places that call it. But I guess my invitation is like, what if we rethought what the unit meant, right? Is the unit a physical, like an actual piece of code? Or is the unit really a concept, an idea, a functionality? So what if we did something like this?
What if we were like laser focused on you know inputs like the setup, right? Or $10,000 loan with the interest rate and the number of periods, and we're very clear of like what we expect as a result And I'm using PyTest here today, and it's got a nifty thing called parametrize. And that allows us to iterate over a list of properties, values, and in this case, these are functions. And these are setup functions which demonstrate the way we can obtain the schedule in our application. Right? We've got the three cases: the calculator API. the loan admin, and then we have the data that's sent when we make that HTTP request. What if we were to do something like this? Well then we would then um
implement these functions, right? The API, the first one was pretty simple. We just return the data as is. The second one's got a little more setup, right? We have to create the super user, log them in, we've got HTML and then blah blah blah. But we get it into the format that we've come to expect. And then finally our um request out to the internet, right? Well Get what's intercepted from the responses, right? And we'll pluck out the property that we care about, which is the schedule. Right. There's other pieces of data that are being sent, but we're testing that separately. Here we're focusing on the schedule And if we did something like that, that would give us three tests that kind of reuse the same setup and assertion. And that would give us a pass and some pretty good code coverage. Pretty cool. And then of course you might guess it, we haven't handled that number of periods
zero. So let's write another test for that. And we'll enumerate the various ways we can get a schedule in our application. and uh we'll get our full coverage. And so what's cool about this too, I'll just go back to this slide, is that like if we have more instances of the schedule throughout the app, right, we'll just add them to this list and it'll just scale out and we can pound it with all our permutations. So why test this way? Why not test the get amortization schedule function directly? Well In our world, it's really an implementation detail. It's kind of like a private method, right? Users aren't like this isn't an SDK that we're pip installing and users are using, right? It's really an implementation detail that could be refactored, it can be moved, it could be outsourced to another library.
And by testing in this way, we've created a world in which all those possibilities can exist. So let's drill down a little bit. Like, why do we want to refactor? The reasons seem obvious, and being able to refactor gives us like flexibility, speed, productivity, and perhaps that all leads to profit. But I think it's about something a lot more important and a lot more fundamental, almost like something to do with our health and our well-being as developers, right? Writing tests in this way, allowing for future changes, right, is an act of kindness to others, but also to yourself, right? We're striving to create an enjoyable work experience, to have that joy of being able to make improvements and progress, not being stifled when, you know, we
're we're attempting to do that. So my invitation is you can make a difference, right? In your test, like if we focus on our outcomes and uses rather than defending implementation choices, we can create that world. By seeking alternatives to mocks or avoiding them, right? As we've seen today, there are alternatives to it. So there are a lot more out there. And so we kind of want to see if mocks can be the last resort. And then finally, everything we talked about, right? Fighting the implementation bias, these things kind of happen naturally if we write our tests first. Like how can we be biased by an implementation that doesn't yet exist? And so the takeaway. My invitation.
The next time you're sitting down, you've got your laptop, you're in the groove, right? And you come to write a test. Right, ask yourself, right, does this is this test that I'm writing enabling or is it hindering like future changes? Right? Am I testing with empathy towards future developers I'm gonna come back to give blame. This talk was really focused on future proofing, right? Sending the empathy forward to the next person. But put yourself back in 2016 Mark, who's raging, who's angry, who wants to blame, literally, um his his anger on the author, right? Do we really think that person woke up that morning and set out to ruin my future day, right?
Is that like a reasonable thing? To think? No. And I think fundamentally, right, people are doing their best, they're trying their best. And perhaps a little bit of empathy in this situation could have also helped, right? We don't know what was going on in that person's life. Right? They might be underslept. They might be up to their neck in diapers and up till 4 a. m. or Maybe the way that this was written was imposed by the culture at the company, right? This is the way we test. Or perhaps this is just the way the person knew how to write tests, right? Maybe this was their first time writing tests. Maybe they've been working on code bases that had no tests and they were coming along.
Setting up a test suite and then writing tests to like defend this implementation to protect it right against bugs and regressions. Right? I ran git blame and then I learned something. It was I who wrote these terrible tests back in 2016 and created this suffering. I have the uh the repo here if I've put all our examples into a repository. There's a lot more to it. So if you want to cruise through that, I'd like to thank the organizers of DjangoCon for putting this event together. That's a huge undertaking. and uh giving me the opportunity to speak today.
This is very special. And I'll share a little personal story with you. Over the last four years I've kind of experienced a lot of hardship. I've been battling Hodgkin 's lymphoma. And uh I started out, had the treatment, was okay, and then I relapsed. And the cancer had gone into my uh spine, and so it was like a stage four relapse. And so after 44 chemotherapy treatments, 15 fractions of radiation, and a stem cell transplant last summer. Um I got uh a Cleedeville of Health last week just before coming here. So wanted to share a feel-good story. Thank you. So thanks for being here today. This has been my first talk at a tech conference ever.
So um very special to be here with you. So thank you.
This happens because of implementation bias: the tests are checking how the code was built rather than what the application does. They can therefore defend a broken implementation and break when an equivalent implementation changes.
Discussed at 6:04It means designing tests with compassion for the person who will later need to change the code, and with humility about the fact that implementations and dependencies will change. Tests should enable future improvements rather than preserve one implementation forever.
Discussed at 7:39Test through the interfaces and behavior a user relies on: make requests to the real URLs and assert the returned data or rendered result. Avoid importing views, calling private methods, and asserting implementation-specific mocks.
Discussed at 9:11Create and log in a superuser, create a real database object, request the exposed admin URL, and assert the response and relevant HTML data. A DOM-querying tool such as PyQuery can make the HTML easier to inspect without coupling the test to the admin implementation.
Discussed at 10:46Use a request-interception library such as `responses` instead of patching the `requests` library at a particular import path. Register the expected URL, run the code, and inspect the intercepted method and payload.
Discussed at 13:52Freeze time with a tool such as Freezegun, then assert the exact timestamp expected in the outgoing data. This avoids fragile comparisons against a clock that changes during the test and does not care which production implementation supplies the time.
Discussed at 15:22Treat the functionality as a user-visible concept rather than a particular function, and parameterize a shared setup and assertion over each access path—for example, an API, the admin, and an outbound notification. Add each new usage to the parameter list and test edge cases, such as zero periods, across all of them.
Discussed at 19:13Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026