Error Culture with Ryan Cheley

This video features Ryan Cheley at DjangoCon US 2024 in Durham, North Carolina, USA.

Error Culture with Ryan Cheley
0:26:33
Published December 6, 2024
170 views

In my talk titled "Error Culture," I take a deep dive into the widespread but often overlooked issue of how organizations manage error alerts in technology and programming domains. The core of my discussion revolves around the concept of 'error culture'—the organizational tendency to ignore or minimize the significance of error notifications due to their frequency and the false perception of irrelevance. This paradox arises from our dependency on technology and the need to be alerted to failures, which results in a deluge of alerts, many of which are false positives or not actionable, leading to a dangerous desensitization to potentially critical warnings.

Through my analysis, I address the serious consequences of error culture, such as the risk it poses by allowing significant alerts to slip through the cracks and the challenges it creates in integrating new team members who lack critical, undocumented knowledge. I pinpoint several root causes behind this phenomenon, including alert fatigue, misunderstandings about the alerts' purpose, and a culture that inadvertently celebrates crisis resolution over preventive measures.

My experience has led me to identify certain symptoms that indicate the presence of an error culture, such as an over reliance on email rules to filter alerts and a widespread uncertainty about why certain alerts are received in the first place. To move forward, I propose a refined strategy for managing alerts to ensure they are meaningful, actionable, and directed at the appropriate recipients. This approach involves clarifying the significance of each alert, tailoring them to the correct audience, and providing actionable resolution steps.

This talk will leave you with a call to action, urging a shift towards a more mindful and efficient approach to error management. By reevaluating our interaction with technology alerts, I believe we can transform them from noise into valuable signals that enhance organizational efficiency, resilience, and growth. Join me as I outline practical steps for combating error culture and fostering a proactive error management environment within our organizations.

This talk was presented at: https://2024.djangocon.us/talks/error-culture/

LINKS:
Follow Ryan Cheley 👇
On Mastodon: https://mastodon.social/@ryancheley
Website: https://www.ryancheley.com/

Follow DjangoCon US 👇
https://fosstodon.org/@djangocon
https://x.com/djangocon

Follow DEFNA 👇
https://www.defna.org/

Video production by Confreaks
Follow Confreaks 👇
https://confreaks.com
https://x.com/confreaks

Summary

Ryan Cheley defines “error culture” as accepting and ignoring automated error notifications until a problem becomes a crisis, creating alert fatigue and reactive firefighting instead of proactive problem solving. He explains that it grows from unclear alerts, default notifications, poor communication, and the appeal of being seen as the hero who fixes a large incident rather than preventing it. His remedy is to question the purpose of every alert and ensure that important alerts are actionable, provide enough context, and reach the people who can act on them. In the questions, he recommends tracking lower-priority performance problems as issues, tuning thresholds, and documenting why noisy security alerts matter and when they indicate a real threat.

Key takeaways

  • Error culture develops when people routinely ignore alerts and wait for problems to become emergencies.
  • Alerts should be removed at the source when they are not important, rather than hidden with inbox rules.
  • A useful alert identifies the affected system, explains why it matters, and includes a clear action or verb.
  • The right recipients are people who can actually perform the required response, not merely people affected by the incident.
  • For noisy but meaningful alerts, use issue tracking, thresholds, and clear documentation to distinguish normal events from serious problems.

Summarised automatically from the transcript.

Transcript

4,017 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:20

Speaker 1: Well yeah, my name is Ryan Sheeley. Uh I am the senior director, senior regional director of business informatics at my day job. That's just really a fancy title for Director of Engineering. I was also a Jenga Not space navigator. What what Um thank you. Uh I'm also a husband. My wife and I just celebrated 20 years of marriage earlier this summer. Thank you again. And I'm a father who has successfully sent his daughter off to college. Dropped her off about a month ago. So lots of great stuff happening. Uh I'm also a huge sports fan. Uh love baseball, specifically the Dodgers. Yeah, yeah, I know you're here I'm also from the Coachella Valley where we have a hockey team.

1:06

Speaker 1: It is the desert. Hockey in the desert is a real thing. I love me some hockey too. I have a slide where you can find me. There'll be uh the slide again at the end with some QR codes, uh, so it'll be a little bit easier for you to find me. Uh but today I'm going to talk about error culture. And if this talk had a subtitle, which does, it would be, what the heck are all these emails I get for anyway? But first a definition. I want to define what an alert is. And an alert is a warning of a danger threat or a problem, typically, in most instances. With the intention of having it avoided or dealt with, which may be a little bit different than the alerts or warnings that we get in tech, but also in other areas of life.

1:57

Speaker 1: And I'm gonna make some assumptions. The assumptions are that the alerts that you get are done via email and that they're automated. And this is really a conversation starter. I'm going to have a spot at the end for questions where I'd love to find out what your questions are. in the hallways during Django Khan. But what I'm really hoping is this sparks a conversation for when you get back to work later this week or next week. About how you can better eat uh deal with these alerts that you're getting with being in an era culture I'm going to talk about what agriculture is. I'm going to talk about why I think it happens. I'm going to talk about when it starts, who it happens to, because it's not just us.

2:44

Speaker 1: I'm gonna talk about how you can tell if you're in it, and then if you are, how you can get out of it Okay, great. So I've talked about agriculture here a bit. Has anyone ever heard the term agriculture before? Alright, great, then you're all in the right room. So yeah. I define error culture as a culture that accepts error notifications and ignores them. Encouraging a reactive firefighting culture instead of a proactive culture of problem solving. And you might rightfully ask, well, is that so bad? Like it works, whatever. And you may know, you may be able to guess what my answer is.

3:32

Speaker 1: I'm up here talking about it. I'm gonna say, yes, I do think this is bad. But why is it bad? Well, it encourages a low signal-to-noise ratio in terms of the emails and alerts that you do get. It also kind of helps to encourage waiting until it hits the fan to really solve a problem, which is probably not the best way to approach things. But why does it happen? It's not just like it came down from the heavens and was given to us, although maybe it kind of was. But I think it happens fundamentally because of a lack of understanding. A lack of understanding about what is the alert? A lack of understanding about why the alert is important, and a lack of understanding, most importantly, about who

4:18

Speaker 1: it impacts. But why does this happen? Well, it can happen because of error or alert fatigue. You or maybe some of your users may have reported something very similar to you that looks like this on the screen. Where if I have to click OK one more time But I also think it can happen Because of hero culture, and this might be the most insidious cause of it. These next few slides come from one of my favorite comics, The Work Chronicles. This particular one is Prevention and Cure. And in the first pane, what we see are two people with two fire extinguishers and two small fires that represent problems.

5:10

Speaker 1: In the next pane, we see our hero watching their fire, their problem get bigger and bigger. Until it's so big that they stop and say, hey everyone, we have a big problem. They take care of the problem by putting out the fire while everyone watches, and then of course are proclaimed a hero. How many of you have ever been the person on the left? It's okay to raise your hand, it totally is. How many of you have ever been the person on the right?

5:57

Speaker 1: Thank you. Which one feels better? The one on the right. Right? You get recognized for having done a thing. You're the hero. Everyone knows your name now. Which one is probably better for problem solving, security, user satisfaction? just general maybe mental health satisfaction as well probably the one on the left Okay, but when when does it start? I think it starts for two main reasons, which this is kind of a you know very broad generalization, but internal reasons and external reasons. The internal reasons

6:42

Speaker 1: Or things like someone might say, we collectively need to be notified when this happens again. Or someone says, it might be useful to know about that. I don't really know why, but it might be useful. Sometimes you just get opted into these emails or these alerts. No one asks you if you want to get them. You come in one day one Monday morning and suddenly your inbox is flooded with alerts and you have no idea why. From an external perspective, it can happen because you may bring in some consultants and they they may say, well, this is best practice. They don't provide more context or even really steps on what to do. They just might say it's best practice, and you accept that.

7:28

Speaker 1: And that seems reasonable. You can also get software that's procured externally, installed, and has just default alerts enabled. And you get the alerts, and again, there's no context that's associated with the alert to just come in one night Monday morning because the system crashed over the weekend, and okay, there it is. Well, who does it happen to? Well, I mean, we're at a tech conference. Obviously, I'm going to mention folks in tech, people like developers, many of whom are in this room, but help desk folks. System administrators, network administrators, directors of engineering. Hi everybody, chief technical officers

8:15

Speaker 1: But it can also happen to office workers, administrative assistants, office managers, customer service reps, account managers. And it happens in lots of sectors. I work in healthcare, but I have friends in education that I got to present this to, and she said, oh yeah, all the time. I'm sure it happens in agriculture. We're staying at a hotel, so I'm going to say hospitality, it probably happens too. So it can happen to anyone. Okay, great. So I've talked about what it is, I've talked about who it impacts. Well, how can I tell if I'm in it? I think ask yourself a few questions. If this looks familiar to you, this is a

9:02

Speaker 1: uh uh the the the trash of uh of an email inbox With a whole bunch of unread emails from a no-reply style email address. So is your deleted items filled with lots of emails from a no-reply style email address that you didn't even read? You just deleted them. Here we have uh an image of a rule at my work we use Microsoft. This is a Microsoft Outlook. uh rule set up that actually just goes ahead and deletes it for me. So question number two, we're all smart people in this room. We know about the rules engines in our email clients. Do we just set up

9:48

Speaker 1: uh uh rules to delete those emails for us so we don't have to think about them. Yeah? Yeah me too. That was actually mine. Maybe you get alerts with no context, with no real information. This mentions that a client library has failed, but like, okay, which one? It mentions an IP address, but good luck finding out which one. I mean, okay. So question number three. Do you get alerts and have no idea what they are, no idea what to do with them? All right. Going back to the work chronicles and prevention and cure, we have our hero being proclaimed.

10:35

Speaker 1: Do you see people who knew about a problem that you both knew were coming, but they waited until it was so big to let everyone know about it and then took care of the problem? Maybe that person was you. I'm not going to pretend like I've never done that, because I have. If you answered yes, To one or maybe all of those questions today or over the course of your career, I would say that you're in an era culture. Okay, so hopefully I have convinced you, at least maybe partway, that era culture is bad.

11:20

Speaker 1: But now what, Ryan? What can I do to fix it? Well, there's good news. It doesn't really matter how big the organization you are in is or where you're at Within the ladder of that organization, whether it's an individual contributor or the chief technical officer, you can make a change. You can have agency over this problem. But where do you start? I'm gonna get some water. Well, a word of caution. Change should not be made until the reasoning behind the current state of affairs is fully understood.

12:08

Speaker 1: And I think that this is best exemplified by the idea of Chesterton's fence, which says Don't take down a fence unless you know why it was put up. In this image what we see are two people on one side of a fence with one of them proclaiming, What a dumb fence. What were they thinking? On the other side of the fence, a little bit in the distance, is a rhinoceros. And the only thing protecting those two people from that rhinoceros is the fence. Don't take down a fence unless you know why it was put up. And how can we find out why offense was put up? How can we find out why the current state of affairs is the way that it is?

12:59

Speaker 1: By asking questions. I'm gonna give you a couple of ideas of what I think are important questions to ask in terms of alerts, but asking questions I think is generally a good idea whenever you're coming across something that you don't fully understand. And this can be super hard. If you've been at an organization for 15 years, you may feel like I should already know this. Ask the question You might be brand new and say, well, I'm worried about asking, and I'm brand new. You are in the best position to ask the question Because you haven't been inoculated against all of the weirdness, all of the things that don't make sense. So ask questions. The first question is, is the alert important?

13:47

Speaker 1: The answer is no. Well, delete the alert. But not like we did before. Don't just delete the email that you got or set up a rule so you never have to read to to read it. And the reason that's important is because you know about it, but the new dev who's gonna start tomorrow, next week, next month, next year. They don't know. You want to delete the mechanism that generates the alert. But if you ask the question, is the alert important? and you get back a yes, Then this is fantastic news because you now have an important alert. Next, ask the question: is the alert actionable? And what does an actionable alert look

14:33

Speaker 1: like? It has a verb. This is the superhero verb from Schoolhouse Rocks. Anyone know Schoolhouse Rocks All right. Yeah. For those of you that don't know, it was an interstitial cartoon in between other cartoons on Saturday mornings. Taught all of us that watched it all sorts of great things. Most importantly, what a verb was I'm going to take you through a couple of examples starting with what I would call a bad alert on up to a best alert, all about the same thing. So, in the bad example, we have our alert. It says, super important alert about the server. And the message is: the server is unresponsive. The

15:18

Speaker 1: server is unresponsive. If you're lucky enough or unlucky enough to work in an environment where there's exactly one server, then this will be helpful to you. But if you don't, good luck figuring out which server it was. So let's update it. Let's put the name of the server in there. Same subject. The server DOWeb005 is unresponsive. Well fantastic, I know which server it is, but I don't remember what to do about it Okay, so we're gonna update it to include a verb. The server is unresponsive to resolve this do X. Do X is the verb here. It might be to reboot the server.

16:03

Speaker 1: So this is great. We have an actionable alert. But do we know why the alert exists? Do we know why it 's important? Not yet. So going back to our best example of an alert, we can update it to include a little bit of context. See this link for details on the alert. And if we click on that link that goes to our knowledge management system, we might get something like this. The server Dio Web005 is a test server on DigitalOcean. It's used for Project ABC, which is set to be retired on September 1st, 2024. If I'm doing my math right, today is September 23rd, 2024, which means that this alert might not be port

16:50

Speaker 1: important anymore. The server might be unresponsive because someone finally had a chance to decommission it. Now, before you go and just delete this alert and the mechanism that generated it, you want to double-check that. But that's probably what's happening here. You don't need to stop everything you're doing. to get this taken care of. However, where I work, if I clicked on that link and got something like this, I work in healthcare, If it said the server is a production server on DigitalOcean, it's a mission critical server for claims adjudication. I would literally stop everything I was doing right then and there to go figure out how to get this taken care of and get it done. I'd go Get the the server rebooted. Now in the alert context or I said to add some alert context here and I had links that went out to my knowledge management system.

17:38

Speaker 1: Where I work, that works for us. But you can also embed the context in the alert if that works for you. There's no one size fits all here for this, but you want to make sure that the rationale Of the importance of the alert is there and easily retrieved. Okay, so who should we notify? Going back to our best example If we click on the knowledge management link and it tells us it's the mission critical server for claims adjudication, we could ask the question: are these the right people to notify the claims team? A business analyst, a developer. It might be good for the claims team and the business analyst

18:25

Speaker 1: to know that this had happened, but they can't reboot the server. It might be important for the developer to know about this, but where I work, we can't reboot anything. So these aren't going to be the right people to do the action and the alert. You want to send it to a server administrator because they're the one or s yeah, a server administrator, because they're the one that can actually do something about it I'm a visual learner. I like to see things in pictures. So I'm gonna talk about what makes a best alert in a Venn diagram. Alerts should be act uh actionable, they should be important, and they should be sent to the right people. If an alert is actionable and important, but not sent to the right people, it just leads to frustration. I know this is important. I know what to do, but I can't do it.

19:12

Speaker 1: If it's actionable and sent to the right people, well it's a time waste. I just spent three hours working on this and you're telling me it's not important? I'm just gonna ignore all of my alerts. If it's important and sent to the right people but not actionable, well I know I need to do something about it. I know I'm the person that's supposed to do something about it. I can't find anything in my knowledge management system about it. Hey Carlton, do you remember what we did last time? Because I I don't. So the best alerts are gonna be actionable, they're gonna be important, and they're gonna be sent to the right people. Error culture can be pervasive. Error culture is pervasive. But we can make it better

19:57

Speaker 1: by asking questions and making sure that our alerts are actionable, important, and sent to the right people. I'd like to give a special thanks to a couple of people who motivated me and encouraged me to be on the stage today. If it was not for them, I literally would not be here. Adrian Frankie, Carolyn Zimmerman, my friend Mario Munoz, and Trey Hunter. And now it's time for questions, and this is where you can find me on the internet. Thank you so much.

20:34

Speaker 2: Thanks Ryan, that was super. Um I've got an alert that I get from Sentry and it comes in every so often telling me about a performance problem which I know I need to resolve and I'm I I like the alert because I know if it's got worse or it's about the same as it was, but it hasn't yet bubbled up to the top of the priority list. It's on the list But it's not got to the top yet. And the rubber the alerts kind of like my nice keep an eye on that. I've got a nice little campfire that I'm warming my hands next to, so to speak. Do you have any advice for that? So how can I improve my flow there? Because it does there is a bit of noise.

21:09

Speaker 1: Yeah. That's a great question. Thank you. I would say that depending on like kind of if it's single person or multi person team or like big organization. I'd actually write up a an issue like a GitHub issue or something like that to so that everyone knows about it and put it there because Okay, yeah, it's a thing, but if you're just going to ignore it, like maybe change the threshold on it so that it notifies you at a different sort of cadence or or or whatever, but Yeah, I would turn it into an issue and and make it so that it got sent to you when it was really a problem versus just an annoyance that maybe someone hit something and it took three seconds versus two seconds to load a page or something.

21:55

Speaker 1: Something like that. Yeah, thank you.

21:57

Speaker 3: Thanks, Ryan. Um so there's there's a kind of alert in security that is usually spurious but sometimes really scary. Like an example might be um in the new in the new Mac OS there's a new like Application X has permission to record your screen. Do you want to continue? That Apple has decided to show you once a month for every application that can do it. And like It leads to a lot of alert fatigue, a lot of people are really angry about this alert, but like I've been guilty of building alerts like that that most of the time are totally spurious, but when they're not spurious. they mean something really significant. I I wonder if you have any advice on sort of how to deal with, you know , s situations

22:42

Speaker 3: where Yeah, there's this tension between wanting to let people know that 90% of the time this is fine, but 10% of the time it's really bad.

22:51

Speaker 1: Yeah. I would I would say that making sure the context, like in that case I would probably link to a knowledge management system on the alert to kind of provide more context for, well, you're going to get this alert. It's going to happen 90% of the time. This is when you can tell it's really a problem. It doesn't necessarily take care of the fatigue part where people like keep getting it and they may ignore it to a point that then suddenly it becomes a really big problem. And like, well, in the Mac OS instance, like maybe they've got bad malware or someone has taken over the computer or something like that. But like having context there, I I think can help with those things. And then just having honest conversations with your teams about like, well, why do we get this alert? Why can't we just turn it off? And making sure that people understand that and that it's embedded in the context. I think so much of this is just solved by better communication.

23:39

Speaker 1: Because like I'm guilty of this. I just want to cool build cool stuff. I don't necessarily want to tell people like why it's you know like what the different things are. I just want to build cool stuff. But but yeah, I would say that adding some some context to it to just let people kind of have a better sense of why it's there and why they shouldn't always ignore it. But that's a super tough one. Oh, we got a question back there.

24:05

Speaker 4: Thanks. Is this on? Yeah, thank you very much. Do you think there are some insights and lessons from the specific case of alerts and the way people respond to them? that can be applied, for example, to documentation and communication in general.

24:27

Speaker 1: I'm s I'm not sure I totally so alerts can you ask the question again?

24:32

Speaker 4: You identified some really specific uh problem patterns that occur with alerts and uh the way people behave in in response to them um perhaps as a result of actions that were missing from the alert and so on. Now maybe this is a specific uh case alerts, but maybe the what your the insights you have can be applied to other to documentation, for example.

25:05

Speaker 1: Yeah, I mean I think in general all of our documentation could be better in all sorts of different places and ways, and making the user interfaces that are that that we have designed easier to use, easier to understand, easier to follow so that the challenges of the users aren't so So challenging. D does that answer your question? I feel like I'm not I didn't select on anything. Oh, okay. Um

25:46

Speaker 5: In the interest of time.

25:47

Speaker 1: Oh, oh, I'm being being kicked off the stage.

25:50

Speaker 5: No, uh, Ren will be available for questions in the hallway, so you can reach out to him. Yeah, so thank you for coming. Thank you, Rain. And the lightning talks in the junior boardroom. So if you need to attend some, you're welcome. And thank you.

26:02

Speaker 1: Thank you very much, everyone.

Questions this talk answers

What is error culture?

Error culture is when an organization accepts and ignores error notifications, creating a reactive firefighting culture instead of proactively solving problems. It leads to alert noise and delays action until problems become serious.

Discussed at 2:44

Why do organizations end up with too many ignored alerts?

The main cause is a lack of understanding about what an alert means, why it matters, and who it affects. Alert fatigue, automatic subscriptions, consultants’ supposed best practices, and software with default alerts can all contribute.

Discussed at 3:32

How can I tell whether my team is in an error culture?

Warning signs include deleting unread no-reply alerts, creating email rules to hide them, receiving alerts with no context or instructions, and waiting to report known problems until they become emergencies.

Discussed at 9:02

How do you fix error culture and improve alerts?

First understand why the current process exists rather than removing it blindly. Then ask whether each alert is important and actionable, include a clear action and useful context, and send it to the people who can actually respond.

Discussed at 11:20

What makes a good or best-practice alert?

A good alert is important, actionable, and sent to the right people. It should identify the specific system, say what action to take, and provide enough context—directly or through documentation—to explain why the alert matters.

Discussed at 16:18

How should I handle a recurring alert that matters but is not yet a high priority?

Track it as an issue, such as a GitHub issue, so the team knows about it, and adjust the threshold or notification cadence so it alerts people when it becomes a real problem rather than a minor annoyance.

Discussed at 21:09

How do you reduce alert fatigue when an alert is usually harmless but occasionally serious?

Add context explaining why the alert exists, how often it is expected, and how to recognize the genuinely dangerous cases. Discuss with the team why it cannot simply be turned off and make that rationale easy to find.

Discussed at 22:51

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Ryan Cheley

More videos from DjangoCon US