Mixing reliability with Celery for delicious async tasks
Published November 22, 2023
This video features Flávio Juvenal at DjangoCon US 2017 in Spokane, Washington, USA.
DjangoCon US 2017 - Preventing headaches with linters and automated checks by Flávio Junior
While it’s very common to enforce PEP8 code style with tools like flake8, it’s rare for Django projects to use any other types of tools for automated checks. However, linters and automated checks are a good way to enforce code quality beyond code style. Human-based code reviews are great, but if an experienced programmer leaves the organization, all quality-related knowledge they have will be gone. One way to prevent this is to make developers consolidate their knowledge as custom check tools. Instead of repeating to every junior programmer how they should code, experienced developers should write tools to do that for them. Having this kind of “executable knowledge” is great to ensure long-lasting good practices in organizations.
Thankfully, Python already has a number of extensible linters and check tools that can be used to consolidate knowledge. Also, Django has the System check framework, which can be used to write custom static validations to Django projects. In this talk, we’ll discuss existing linters and tools, what benefit they bring to Django projects, how to extend them and how to build custom ones. Combined with IDEs, pre-commit hooks, and CI tools, linters can validate code at programming time, commit time, or CI time, thereby ensuring good practices in all development workflow.
This talk was presented at: https://2017.djangocon.us/talks/preventing-headaches-with-linters-and-automated-checks/
LINKS:
Follow Flávio Junior 👇
On Twitter: https://twitter.com/flaviojuvenal
Official homepage: https://www.vinta.com.br
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Linters can do more than enforce style: they can catch common Django bugs, insecure patterns, documentation pitfalls, and project-specific rules before they reach production. Flávio Junior explains Django’s dynamic system checks and several kinds of static analysis—text, token, AST, and inference-based—and shows how to write checks for shared mutable `ArrayField` defaults and forgotten `QuerySet` returns. He recommends running checks in editors, pre-commit hooks, CI, and code review, and argues that teams should turn repeated code-review knowledge into automated checks that preserve good practices and help train developers.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Thank you all for attending. Is it fine Oh. Hello, hello. Okay. Uh if any of you can't hear me well, just let me know. And okay, so the slides are here. Uh really happy to be with you all today here. Uh I'm finding this conference really great. Really congratulate the organizers, the sponsors, everyone who is volunteering.
Speaker 1: It's been really nice. Uh if you want to check the slides, they are here. Uh I work at Vinta uh from Brazil uh on a city called Recife, the city I was born in. And I decided to start Vinta because I couldn't find a great place to work with Django and Python on my CD and started Vinta with other partners and we work with companies companies worldwide uh doing development services, custom web development with Django and React. And I'm going to talk about linters, what are they, uh why using them, how to implement them, uh when to run them, and which ones exist and we can use Using our Django projects. First of all, what?
Speaker 1: What are these themes? Uh I I'm not sure if uh Native speakers do know that. I I didn't knew, but LinkedIn is that little thing. That little thing that's called I I really like that definition clinging fuzzy fluff that accumulates. And this last part, and we hate, I just added it, but you can check that on Wiktionary. Uh it's a real definition and also uh good suggestion here, clinging fuzzy fluff, I think it's a really good name for a indie band. So if you are looking for a name for your indie pen, that's an idea here. And a linter uh it's not Not many people call this a linter, but some do.
Speaker 1: It's a thing that removes the the little stuff that we hate. And this is also valid for software. There's linters on software that analyze code to find flaws and errors, helping to remove that clinging fuzzy fluff that we rate. And here is a very basic example using Pyflakes, I think the most well-known linter on Python, that it can detect that your my sum function uh uses it is duplicating the argument text and it it can find that and this is quite good because it's a a common bug that we may uh we may do during coding Okay, but why? You may not be convinced, okay, uh, but I I don't make this kind of mistakes
Speaker 1: like I I'm expert developer. I don't need that. Okay, uh what's wrong with this code? Can anybody uh find the error here? Can anyone what's wrong Ten seconds. This is a return, yeah. This is in the return. Like if you are overriding query set, uh remember that filter exclude all those things return a new query set. So you need to attribute it again. to the variable you are working with or you need to return. So this is a subtle bug that many Django developers already did. And can our
Speaker 1: linter really detect this kind of things? Yes it can. Yes they can and we'll see how. And linters prevent bad style. It's the first thing that linters were made for, but they also can prevent bad patterns. uh bugs and even security vulnerabilities. There are linters for that in Python. And another thing that many people don't consider but I think it's really important is that expert developers know language framework library related quirks and common mistakes And maybe those expert developers should write that knowledge up down as linter checks. And this way the the knowledge is perpetuated into shacks and it will go and do a lot of benefit to the
Speaker 1: the community. So for organizations linters can consolidate knowledge in their forum And this way you can enforce long-lasting good practices, you can automate code quality checks, and you can even help to train new developers. And in fact, organizations are already doing this. Uh I don't know much about closed source uh organizations, but open source orgs are doing this. Uh Twisted has its own checker with uh piling it and flake it plugins, OpenStack 2, EDX2, SouthStack 2, all of them implemented with PyCode style, Flicate, Pylene, Bandit plugins, because those linters are extensible and you can write your own checks for them. And the kind of things that organizations are checking are like if you are using some blacklisted module, some module that should be test only and should should
Speaker 1: be used in the real uh code the f for the project on the project. In consistent string formatting unfortunately uh string formatting in Python doesn't respect the single way single obvious way to do something right now there are many ways to do string formatting Python so you can write a check to prevent that. You can like make people aware that they are forgetting to call Supper on unit test setup or Town. And anyway, there are many many things you can check and organizations are doing this. Also another thing that linkers do is that they make better UX for libraries and frameworks Because error prevention is good UX. And in fact Django does that. Many people don't know, but Django comes built
Speaker 1: in with some checks. the system check framework with some built-in checks that you can run with python minage. poi check. If you have a pr for example example project That's duplicating the namespace for the URL into two different apps, app one, app one, but it should be app one, app two. And if you run the check, it will detect that for you. to show a warning and it will uh return that command with a non -zero exit uh uh code and then you You can run this even as a part of your CI to check for this kind of stuff. Uh but there are many checks that Django doesn't implement yet, but it could. For example
Speaker 1: So if you if you use an array field, this is straight for Django docs, uh you should make sure that you are passing a callable uh default Because incorrectly using default with a empty list creates a multiple default that is shared between all instances of a ray field. And you probably don't want that And a suggestion for Django, and we can work on that on springs. Uh I'm really uh happy to join someone to work on this, is that perhaps every insure remembers to don't forget or similar warnings that are on the docs should become a new system check and this way uh developers will be able to check their own code and prevent bugs or mistakes Because they forgot to read the documentation well
Speaker 1: or they missed something. Okay, but how? How can we write those checks? Uh we can either do via dynamic analysis or static analysis Dynamic analysis is performed by executing code, not necessarily at runtime. You can just run the code to check it, but you have to run the code. In order to work, certain checks need need to be dynamic. For example, they need to execute Django, they need to introspect models or something. Uh a check for unapplied migration is an example. You can't do that uh statically. You need to connect to your database to check if the migration is Applied and Django System Check Framework is dynamic. Uh we can solve this problem that I just told you
Speaker 1: of setting a Empty list default to array field with a dynamic check. Let's see how. So live demo because live demos never go wrong Uh so I have here the model, okay, with empty list as the default and it It's an array field, so it's wrong because other uh like outposts which are the same default, the same tags It's not good. And we can write uh like a checks. py file or something uh that we import the register from checks. We register the function, the check function And inside this check function we iterate over all models from all apps
Speaker 1: And iterate over all fields for these models. And if the field is an array field, if the field has a default And this default is not callable, then this is an error. And we just append the error here in this errors list. And the error says Field uses a default instance that's shared between blah blah blah use a callable instead and then add this error to the error list return the zero if we run like I I had to Of course, I had to add this checks app here on styled apps. Okay, so just added the the app here. And if I run manage. py
Speaker 1: Check it says uh that it found that app post. tags uses uh uh default instance that share between now field instances and should use a callable instead. And if we fix here and run it again The the error is gone. So we wrote a check that really works and is really doing the right thing, preventing bugs That's great. But there is also static analysis and that's performed without actually executing the code. And because of that it's safer and more general than dynamic analysis It can analyze how code flows.
Speaker 1: So when dynamic when dynamic analysis goes inside AIF, it goes obviously into one of the conditions because it's actually executing the code. code. With static analysis, you are looking through the code, not executing it, so you can analyze uh all branches, all code flows. It's similar to code review, but it's performed by machines There are many types of static analysis, mainly for. This is something that I divided, like not academics divided like that, but there are text. and regex based static analysis there is token based static analysis, AST based and inference based. That's and rejects based it 's ideal for simple checks you are just like
Speaker 1: looking for something on your code like just a simple contains check or a simple rejects mat match Check. And there is a library that does something like that. It's called DOGI. And it looks for Python at Python code, searching for things like passwords or diffs that someone forgot that. There are there is also the token-based approach which uses the tokens that the language has to analyze the code. Here's an example. You have this file. Note there is a space between the prints. Print function and the open parenthesis. And if you tokenize this little file, you see that the tokens are name, print, and operator,
Speaker 1: parentheses And there is a distance of one between them. So you can write style checks with token-based static analysis. And in fact, PyCode style uses this So it's better to get structure than raw text, obviously. It does not lose info. So if you tokenize some tokenize a code and untokenize it, you get the same code back It's ideal for style checks and Pycode style part of play-cate uses this. And more mo more interesting, there is the abstract syntax three based checks. Uh here's a example of an abstract syntax tree. You can run that code if you want. If you just install the Stor
Speaker 1: library to pretty print the AST but you can just parse the AST. AST is part of uh Python built in modules and you can dump it with Astor and you See something like that. That's a tree representation of the code. The code is a module, it has a function definition inside it with arguments and a body that has a return. That's an operation between left and right operators, and it's an add operation and everything. So it's a tree representation of the code And it abstracts some if or away. So for example, if else elif else becomes nested if else. So that's why it's called abstract because it abstracts some if or away from the
Speaker 1: code. And it's ideal for checks that need to analyze the structure of the code as a whole. Checking the relationship between parts, for example, logic errors like undefined name And Flakeate uh uses AST based checks, like on like you can see in this example up here And the ST is made for walking, okay. Uh you you can you can you can walk through the AST Using a node visitor abstraction. And here is a just a simple example of uh some code that prints all the functions inside the inside uh AST. So it's it's How can I say it's uh recursive okay
Speaker 1: because after visiting the function definition you keep visiting uh the other stuff so it it you have to call generic visit again. Uh but this This little code here just prints all the all the functions it can find on the on the code. And we can solve that problem we solve that we forgot to return on the query set with AST walking. We just need to know that we must We must check for calls, okay, uh because we are we are checking for a call to filter and we must check for expression node Because expression node is like a function cow or a stra expression something that's not returned or stored in a variable or something
Speaker 1: Like that. So it's like a function call that does not return or is not assigned. And exactly that what we are looking for. We are looking for a call to filter that's an expression And demo for that? We have here the code to check to find this bug, okay? To find that we are missing the return here. We need to define a visitor. And this visitor visits expressions. That's what we are looking for. And this expression needs to have a cow inside it. It needs to have uh the the called function must be an attribute of a class
Speaker 1: because it's self dot filter This method needs to return a query set. Obviously, we need to teach that to the code, so we just create here a list of querieset returning methods, filter xclud out. Is larger than that, just an example here. And these methods need to be called over a name, and this name must be self because we we are looking for self. filter. So you this This part here checks that it's an attribute, dot filter, and this part below checks it's self dot. And if If we if all of that is true, we just found uh self dot filter uh and as an expression it's missing The return probably. And then we can just print
Speaker 1: query set expressions, not assign it. This is an error. One thing that's missing here is that I'm visiting expression on all the code, the entire code on all classes. I need to make sure that I will only visit if it's inside a class that inherits from query set. But this is very easy. I just need to create a query set class dev visitor. That checks if the base is an attribute and if the this attribute is query set. If I do that, I'm checking that I'm inside a class that's inheriting from models. query set and if I if I run this code , uh
Speaker 1: If I run this code, I find that it it finds the the self dot filter thing. And if I fix it with return and run it again, the error goes away. So it's working. That's great. We made a new check, uh a new linter here, but it's not perfect if we If we do something like this, we can't detect because we are detecting self-dot filter. We are not detecting something that returns a query set And operating over acquires that again returns acquire set. We are not checking that obviously. It's quite difficult to check this So, how how can we fix this? It would be great if we could infer the type of self.
Speaker 1: autor scal. And we can with inference-based checks. Uh we can infer info from AST nodes. We can infer things like imports, variables, attributes resolutions, operation results, method resolution orders, etc. We can do something like an interpreter that doesn't actually executes the code but analyzes it inferring some info. And PyLint does that with the this library called Asteroid And we can show that with here the same code, okay, uh the same bug we saw. Uh if we had self. autor scal here, the other linter could detect it, couldn't detect this. But using asteroid that the leap that powers
Speaker 1: violent, we can visit expressions again, okay? And If the expression has a call inside it, and if we infer the result of this call, and the result of this call is a query set. We are we are done. We already found something that returns query set and is an expression. And it's probably wrong because you don't do operations on query sets and do nothing after it you usually return or assign to something. So just by doing that, but by inferring, we found the bug. And if we run this code here You see that it takes a lot of time to infer stuff, but it finds the bug.
Speaker 1: And I printed here what was inferred, and it was inferred that the result, the result of the operation. should be a person query set and that's great but it I'm doing something here that I'm not showing to you I've shown right now I'm teaching teaching asteroid to infer query sets because it queries set code is quite complex. Asteroid can't analyze it and figure out that filter returns a query set of the class that uses the model and everything. I need to teach that to to asteroid but it's not uh very difficult I just need to do the same I need to say that query set returns methods Returns quer sets again. So I just uh say that hey, you have a query set
Speaker 1: instance okay? And if it's uh If this instance is a query is really a query set, for these kind of methods, filter return all, you will infer another query set. Because operations of a query set like filter return another query set. So I need to teach that to Asteroid, but it doesn't take much. Just this code here. And now Asteroid knows how to infer query set That's amazing. Uh but some of you might think what about my pi? MyPy can do type inference. What about it? Unfortunately it's not there yet. I asked this on my pi um my pie issue and Guido answer it really happy with this but
Speaker 1: he said that it would be very useful and powerful to use my pi for inference But it will also be very complicated and this is not something they are working on right now. But I was on Pi Bay last week and there are people from MyPy there and they are really excited about this. So they work on this like sometime, not right now. But we can disobey Guido and just try to infer it. And there is some code for that. It's not very useful you can check there. Uh there are some this code is commented. You can check what I'm doing. But it's not very useful. There are other types of libraries that also do type inference, PyType, PySharm , both make use of type annotations are great uh and have inference capabilities.
Speaker 1: Asteroid don't use uh type annotations yet but it it will use uh and but PyType unfortunately uh has Only initial Py Python 3 support, even though it's used on more than 500 internal Google projects, and PySharm is implemented in Java. Okay, when when to run this? We can run linkers on programming time, commit time, continuous integration time, or code review time. Programming time, most of you must have uh have uh editors that support linters. But there is also some stuff that should never be committed. This is from a real government website from Brazil, you can see here Someone committed a mesh conflict.
Speaker 1: And to prevent committing this kind of stuff, you should run checks on commit time. Okay. You have pre -commit for that. It's a great tool written in Python So if you don't use it, try to use it, it checks for these kind of things too, like conflicts and uh security related stuff or even any liter you want just need to configure continuous integration time. Uh many people don't do that, but you should fail your build if you if any linter reports any issue. It's really important, especially for things related to security. There are linters related to security like bandit and safety, and you should fail your build if any of those links fail. And also code review time. You can have a bot that comments on
Speaker 1: On pull requests, especially useful for open source projects, there's a project from Lyft that does that. Which linters there are out there? Many, many dozens of linters are available. Linters for quality, imports, docs, security, packagings, even spelling. Someone wrote a linter to check the spell Not of the strings of your code but the code itself the variables uh and you can wrap them all with prospector a python only two but if you have a multilinguage project you can use koala koala. io it's a great project that wraps linters from all languages And let's clean all the things. I created a list of all the linters I could find. It's there, hintasoft slash Python linters and code analysis. There are many, many linters that you can use
Speaker 1: And I have some ideas for new Django checks to write, like checking for string formatting in Raoul-esque circuit. X SQL injections, we don't want that. No true in char field or text field. I think we can work on that on sprints. So if you are interested. interested in writing custom checks let me know look for me on sprints I want to make a new project called Django Bookfinder to have checks for that kind of stuff And finally, some people criticize code analysis because it it gives false positives or it doesn't understand dynamic stuff. But maybe if code analysis isn't understanding your code, maybe your code is like too complex. Maybe your fellow developers won't understand it too.
Speaker 1: So let's try to write simple code, testable code, checkable code that our fellow developers then code analysis. can check too and that's it. Uh let here is my contact info. Feel free to reach me. Let's try to contribute boot on this uh Django Book Finder thing that I'm just starting to build with Vinta and you can check the Python linters and code analysis repository for a full list of all linters I could found There are other talks from Vinta uh available on this on this link. We'll give another talk today about uh salary that Filippi, one of my partners there Uh I think that we that will give it. Thank you very much.
Speaker 2: So I really love the title of your talk. I thought it was funny, um because we've all been there, right? Um or many of us have been there. Um so what was it that inspired you to become so passionate about helping, you know, people come on board with you? with making it easier to to check your code.
Speaker 1: Okay. So on on Vinta we do a lot of code review. Our code is code reviewed So I found myself repeating myself too much like the same mistakes again and again because new new people were joining the company and so I I I thought maybe I should write something to check for this kind of stuff. So for example a simple thing we wrote one of the first checks was to check if someone forgot to run make migrations So this like and and this is difficult to check on pull requests like you you need to look that someone changed the models But didn't include the migration on the pull request. So this is even difficult but something that we are repeating ourselves, hey you forgot to run make migrations So we wrote a check for that.
Speaker 1: And we are trying we are still beginning to do that on Vinta, but we are trying to consolidate the knowledge we have from code reviews and years of experience in the form of linters.
Speaker 2: Thanks, y'all. It's break time.
Linters analyze code to find flaws and errors. They can enforce style, detect bad patterns and bugs, and in some cases identify security vulnerabilities.
Discussed at 2:35Linters automate code-quality checks, preserve expert knowledge as enforceable rules, reinforce good practices, and help train new developers. They can also improve the developer experience by preventing errors early.
Discussed at 4:07Dynamic analysis runs code or inspects a running Django project, which is necessary for checks such as unapplied migrations. Static analysis examines source without executing it, making it safer and able to analyze all code branches.
Discussed at 8:05Register a check function, inspect every model and field, and report an error when an ArrayField has a non-callable default. Using a callable instead prevents all instances from sharing the same default list.
Discussed at 8:52An AST-based check can find calls such as `self.filter()` used as standalone expressions inside QuerySet subclasses, where the result is neither returned nor assigned. Type inference can extend this to chained operations or other expressions that produce QuerySets.
Discussed at 14:59Run them while programming, before commits, in continuous integration, and during code review. Pre-commit checks catch accidental files and other basic problems, while CI should fail for issues such as security warnings.
Discussed at 22:40There are tools for code quality, imports, documentation, security, packaging, spelling, and more. Prospector can combine Python linters, while Koala can combine linters across multiple languages; the speaker also provides a repository listing many tools.
Discussed at 24:11Frequent code reviews revealed that the same mistakes were being repeated, especially by new developers. Vinta began turning that repeated review knowledge into checks, such as detecting when someone changed models without adding a migration.
Discussed at 26:50Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published November 22, 2023
Published November 3, 2022
Published November 8, 2018
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026