Boost Your GitHub DX
Published March 30, 2026
This video features Adam Johnson at DjangoCon Europe 2020 in Online.
DjangoCon Europe 2020 (Virtual)
September 18, 2020 - 17h10 (GMT+1)
"How to Hack a Django Website" by Adam Johnson
Why did Facebook have a public Django-based site that got hacked? What was the flaw discovered in GitHub's password reset mechanism that was also found to affect Django auth? Are your projects vulnerable? I'll walk you through some stories of common web vulnerabilities, and what they mean for Django. I've had the pleasure of working on over 50 Django projects so far, so I've seen some patterns emerge.
Adam Johnson explains four ways Django sites can be compromised. A leaked debug page, signed-cookie sessions using Pickle, and a duplicated secret key enabled remote code execution on a Sentry instance; Unicode case-folding bugs could enable account takeover; and unsafe use of `mark_safe` allowed stored cross-site scripting in the admin. He also shows how user-controlled data can break out of inline JavaScript, recommends Django’s `format_html` and `json_script`, regular upgrades, deployment checks, avoiding Pickle, Content Security Policy, and publishing a `security.txt` file for vulnerability reports.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Hello DjangoCon Europe. It's lovely to join you here in Virtual Porto. I'm Adam Johnson and today I'm going to talk to you about how to hack Django. a Django website. This talk is organized around four stories of sites getting hacked. The first one comes from Facebook. It was a Pickle Remote Code Execution attack. Don't worry about what that means, we'll find out in a little bit. This wasn't Facebook itself, but this was a Django-based site that Facebook was running. It was an instance of Sentry. The second uh story comes to us from GitHub. GitHub also doesn't run Django, but it's based on Ruby on Rails, but the same problem existed inside Django. This was something we'd call a Unicode case collision account takeover.
Speaker 1: even longer term, but don't worry, we'll unpack that when we get to it. The third story comes from somewhere I used to work called YPlan where we had an HTML injection attack in the admin. This is also sometimes known as a cross-site scripting. So this would be my personal story in here. And then the fourth one is from many different sites that I've looked at. a specific form of HTML injection that can occur when you're templating JavaScript. So let's get started. Story one, Facebook. This is from a 2018 blog post by a security researcher called Blacklist who detailed everything in this small attack that led them to collect a bug bounty.
Speaker 1: And they thankfully gave us a lot of details that we can expand upon. The first step that Blacklist took was to scan the Facebook IP range. Since they've got their own data centers, they've bought an IP range to host all their public servers in. This is public knowledge who owns which sets of IPs, so it's quite easy to find out. Which servers to scan if you're trying to look into looking at Facebook. Whilst doing this, they made HTTP requests to all these IPs and found an instance of Sentry running. This is a Django-based bug tracking tool and it looked like Facebook had set it up as a prototype. They thought this was kind of interesting. Sentry
Speaker 1: tends to include information about the applications it's tracking bugs for. So there could be other information to find a vector for attack. Whilst browsing around Sentry, Blackgust found that it was running with debug set to true. That's Django's debug setting And it didn't take long to find a bug in Sentry itself that led to an exception backtrace. This is probably something they could even find from the Sentry bug tracker. And from this debug page that Django outputs, they could read all the settings. So Django, when it debugs the settings, it actually tries and filters out the sensitive ones. but uh doesn't stop a combination of sensitive uh of non-sensitive data becoming something that an attacker can use.
Speaker 1: In this case, three particular settings combined to allow a remote code execution attack These three settings were the session serializer setting and this was set to Pickle Serializer. So being able to read this was a useful piece of information. The second was that the session engine used was the signed cookies one. Rather than storing the session sessions in the database on the server, they were being sent to the client in a cookie and the client would send back that session data with each request. And the third uh slightly unfortunate step was the debugging of the secret key. Django does filter out the secret key setting. But Sentry provides a way to override it with its own internal setting in Sentry options. And so the same secret key was repeated there.
Speaker 1: And because the secret key is used to encrypt the sessions, this is what uh allowed the attack. So with those three pieces of information, Blacklist was able to run a piece of Python code locally with their own copy of Django to load, change, and save the cookie that they were using. They could call the Django core signing function to load the signed cookie data. And they were able to do this because they had the secret key. And they knew that the six serializer was Pickle Serializer, which was where the main vulnerability came from. They could then change something inside this data dictionary that got loaded, adding data into the session basically.
Speaker 1: And here they've used a custom class called PickleRCE that we'll see in a second. And then they could uh dump it back out into a byte string. Print that to the terminal, copy that into their browser's cookies, send it to the server, and the server would do the same kind of steps on each request to load and check the session data. And the unfortunate things comes from the Pickle RCE class. So Pickle is a serialization uh module built into the standard library, but unfortunately it's not really a pure data module. It actually is code. Something that you may not know if you're using it somewhere inside your system is that it actually runs an interpreter on
Speaker 1: load to uh like a mini coding language inside Python. and you can customize what will run on load. So that Pickle RCE class that uh Blacklist implemented, it looked like this. It had only a reduce method, and this is called by the Pickle uh serialization method to figure out what functions should be imported and called to load this object back on the other end when deserializing. In this case the class returned os. system and sleep30. So it says to ri to retrieve this object, call this function with these arguments. Well, we can see this won't return any useful object, but it will execute a command on the server.
Speaker 1: When we dump this into a series of bytes, which is what made its way into the session, um You can see that the POSIX part of the OS module is being used. There's the system string that refers to loading that system function, and then here's the command to pass to it. So if we call pickle. loads on this byte string, it runs the sleep30 command. You can get pickle to do whatever you like. You could make it Create HTTP requests, delete data, dump the database, and upload it to a server you control, pretty much anything you want. The Pickle module does come with a warning saying it's not secure only on PICL data that you trust. In this case
Speaker 1: Technically, the data was trusted because it was signed with the secret key, but unfortunately the secret key had been given away. So it was like a two-step weakness. After doing all this, Blacklist was made uh able to claim the bounty. Facebook responded very quickly. They took only 18 hours to patch the system and paid out the bounty two hours after that. As far as their bounty program goes, this was a relatively this was a lower tier bounty. If any user data had been at risk, if this wasn't a prototype sentry instance, they would pay more. So what can we learn as uh aspiring Django website hackers from this? Well, there's really
Speaker 1: uh A few things to take away here. First we can try and look for the debug mode being set. An easy way to do that is to try and browse to URLs that you're pretty sure don't exist. And if you get back a 404 page, that would be Django's debug 404 page. Then you might also want to try and force a crash by pushing untested data and then get a 500 which will output all of the settings that you'd like to inspect. It's also worth looking for things that seem to be using pickle. Anywhere pickle is used there will be maybe some kind of in. And you can also build more evil pickle payloads. In the case of this hack, Blacklist only used the sleep31 to
Speaker 1: prove the vulnerability. But you could build more interesting ones that if you don't control when the deserialization happens, they could ping you on your remote server and you could find out when that happened. If you put on our other hats about defending against these attacks, uh well the main lesson here is never deploy a debug true, like this is something that the Django docs highly recommend against, and we also have the system check framework. to guard against it, which you should run on every deployment using manage. py check-deploy. This will actually exit with a failure code if any setting like debug is set when it shouldn't be We should also try and avoid using pickle as much as possible.
Speaker 1: There isn't a safe mode for it. There are attempts out there to tighten it and prevent it from running completely arbitrary commands, but I think there's normally a way around it thanks to the dynamism of Python. And finally, we should look into deprecating the Pickle serializer from Django. There is an open ticket for this, it would just require a little more work to figure out some corner cases in the sessions framework. Now for story two. This one comes from GitHub. And it was summarized nicely by uh the person who found it in a blog post, John Gracie.
Speaker 1: It was um announced at the end of last year. And they found a way to take over someone's account using some features of Unicode. So the first step here was case collisions. And what is a case collision? So in Unicode, all of the letters from all of the scripts that humans use are represented or are planned to be represented. And uppercase and lowercase are not simple concepts, they're not a one-to-one mapping. So You might see on the left here we have two lowercase versions of the string GitHub that would uppercase to the same GitHub uppercase GitHub on the right.
Speaker 1: The first one uses an I without a dot. This is a Turkish letter and the one on the bottom uses the Latin I with a dot that we're used to using in English. Both these strings map to the same uppercase string. So there's some potential for confusion there, especially if the code we're investigating uh does a case insensitive comparison and performs that by uppercasing or uh one string, uppercasing the two strings that it's comparing. In English there aren't that many uh case collisions. This is uh thought to be an exhaustive list of all of the ligatures and letters in um Unicode that would collide with some English letters.
Speaker 1: But we can see there's quite quite a few combinations still. So if we're trying to find a collision with a string that features a double S, an I or an S, a double F and S T, then there are these letters and ligatures on the left that would collide, the dotless I from Turkish, the Schaufus S from German, and so on. Having found one of these case collisions, um, John Gracie managed to register a domain, which was github. com. without the dot on the eye. All Unicode domains are available for registration now. It depends on the top-level provider as to how much they support, but. com seems to support most of them And they're supported with a system called Punicode, which is a an ASCII-based
Speaker 1: encoding of Unicode, because some of the DNS system does not support actual Unicode underneath So this is the technical domain that was actually registered. And the bug that John Gracie had found was that you could get a password reset for the wrong email address. which allows complete a takeover by just changing the password on that account. So if you went to GitHub's password reset form and you entered John at dotless igithub. com It would search in the database for that string uppercased. It would find John at dottedigithhub. com. and say, okay, that account exists, but then it would reuse the email address input in the form to send the actual password
Speaker 1: reset. So uh you could take over the account of someone uh at github. com or anyone whose email address had a case collision effectively Once you've reset their password, you get in their account, you unleash whatever chaos you want to do. You could export all their data so you can look for it later, delete the account. uh pretend to be them, make uh high profile commits, etc. This bug was in GitHub, which is in Ruby on Rails, but it also existed in Django Contrib auth It was spotted by technical board member Simon Charret after he'd read about uh GitHub's exploit. He checked inside uh Django
Speaker 1: Contribut and there it was. So the fix for this went out on December 18th last year, and it's only in these versions of Django and 3. 1 as well that it is fixed. What lessons can we take away here for hacking Django websites? Well, first we can look for targets with Unicode case collisions. If we're trying to hack a website that has an I in it, then we already know that there's a case collision. If not, perhaps we want to look inside usernames for double S's, I's, and anything else that was in that table. It would also be expanded if you weren't looking in English. There's probably other case collisions in with letters in different languages. It's also worth knowing other freaky features of Unicode that programmers tend not to account for.
Speaker 1: For example, there are a lot of characters that look the same, and these have been used in a number of attacks like phishing domains that seem to be from your bank but aren't actually. And since there this is a known vulnerability that affects uh Django and we could find out and we know which Django versions it affects. If we could detect a site we're trying to attack's Django version, then we could uh figure out whether this attack would work without much effort. There are some tools out there for doing this, called like Django version detection, etc. And the main way they work is by accessing a known file like the admin CSS file. and uh checking for strings in it that were known to be added or deleted in specific Django versions.
Speaker 1: This can normally fingerprint the version of Django quite accurately. If we put on a defense hat, what can we do here? Well, I think the the main lesson to take away is to upgrade Django. This kind of weird, obscure attack that nobody really predicted that affects even the best sites in the world. You can't proactively defend against it, but if you keep on the latest version of Django, you will have all of the uh accumulated wisdom and defense against these. So definitely it's worth subscribing to the Django web blog and keeping track of the latest version. My third story comes courtesy of Y
Speaker 1: Plan, my employer from 2014 to 16 We were an events e-commerce app and we had several of the usual features that you'd expect like accounts, the ability to set your name, purchase things and see your purchase list, etc. My colleague Tom Granger figured out that if he added HTML to his name in the app, it would get rendered in the admin. So his first step was to open the app and set his name to Tom and then a script tag that included alert catface emoji. Then he just needed to wait. Whenever a staff member would browse the admin, then uh his cat emoji face would appear
Speaker 1: and he we would get complaints. Uh what was the problem here? Well the main problem was within our custom admin code, we'd use this mark safe function, which is decidedly not safe. Our user admin class said to list a few fields. One of these was called FullName, which we defined on the admin class as the concatenation of the first name and last name. For some reason we wanted some HTML wrapping that. So we added the HTML inside that function and when we called mark safe on the resulting string. And MarkSafe is needed for this HTML to not get escaped and show as actual literal less than, greater than signs in the HTML.
Speaker 1: Unfortunately, this meant that we were implicitly trusting all of the data that users could set in their first and last names, which meant anyone putting a script tag in there would get that script executed. were from within the admin page. All of the HTML that users provided was basically unsafely injected into the admin. This is what we'd call an HTML injection attack, or it's some kind sometimes called cross-site scripting, because you can include a script from another site. The solution here is to use a helper function from within Django. It's called format HTML. And this can take your trusted piece of HTML with format tags similar to how string.
Speaker 1: format works. and safely add the untrusted user data in at the appropriate points. If we are doing a list of things together, we could also use the format HTML join function. What lessons can we take away as hackers from this? Well, we can try adding HTML to every field that we have control over. It's quite common that there'll be some kind of injection vulnerability, unfortunately And moreover, uh it only needs to be in one field that we have control over for it to work.
Speaker 1: Even if the developer gets 99 out of 100 correct, the one broken one would allow us to execute whatever HTML we like on their website. One way to keep a track of this is to use the beacon resources. If we had script tags with a source or images that referred to different URLs for each field that we put them in. The first one that we see get an actual hit, that could be the one that we know um that lets us know which fields are vulnerable And this is especially important for things like the context we were looking at here where the actual execution of the HTML was in the admin area that we would never be able to access as a pure customer. Putting on our protection hat.
Speaker 1: Well, I I think the first lesson is not to copy any code with mark safe. I think there must be a tutorial out there that we'd copied the admin code from because I've seen the same pattern appear in several other admin classes since. One thing that has been discussed several times on Django to developers is to rename the mark safe function because it isn't safe. And perhaps the name is leading people to think that it is it's safe to use. arbitrarily. And the third defense will be through this security header called content security policy. This ties into my talk last year on security headers And with content security policy, you can tell the browser a policy for blocking resources in scripts and images, etc.
Speaker 1: that you don't trust. So that would allow you to basically ban the inline script tag that Tom had used to alert the CAD emoji and any other sources that you don't trust, and it would never be a problem. And now for our fourth story. This one has happened on lots of sites that I've looked into. As a solo consultant, one of the first things I do on a project is run a short audit. And in pretty much every site I've looked into over the past year, year and a half, I found some variant of this. It depends on, it does vary how exploitable it is, but it's a recurring theme. It's a bit of a variant on the last one. Again, we'd include some HTML in a field that we control.
Speaker 1: and and we 'd see this HTML get injected later. But in this case the key is that um the HTML starts with a closing script tag. And we'll see why this is important in a second Normally there's a view that does something like this. So there'd be some structure of data to be passed through to the JavaScript that runs on the page. And that could be like this dictionary here that includes the user's name. And this is being turned into JSON here, a JSON string, passed through to the template in this user. json variable And then the template inlines the JSON inside a piece of JavaScript. So
Speaker 1: here we have it happening. JSON is a subset of JavaScript, so this looks like it would be safe And that's and the the pipe safe is needed here. Pipe safe is uh that function we saw earlier, mark safe. just by another name, that's the template filter version. Well, you might have guessed that the safe is not safe here, because we would get that cat alert appearing from my name The problem here is that the HTML parses first. So this is the output from that template fragment we were looking at. And at first glance this looks like a script tag, a single script tag
Speaker 1: that contains a valid piece of JavaScript in the middle that sets the user with a name and then there's uh some stuff inside the string but this is all a valid JavaScript string so it should be fine. Unfortunately the first impression is wrong. The browser when it's parsing the HTML does not parse the JavaScript within it. So this first script tag starts a context whilst the HTML parser is running that's simply looking for a closing script tag and so it ends at this point. So this first script tag has basically a fragment of JavaScript. And when this gets executed, it would be a syntax error because there's an unfinished string here. Then there's this second script tag that I'd entered in my name
Speaker 1: and the JavaScript within this is perfectly valid. It's a call to alert a cat face, so that's what we get. And then there's a trailing piece of code with a closing script tag, and I think most browsers would simply ignore this because they try and parse HTML in a very lenient way. So what's the solution here? The solution is a filter that's actually built into Django since version 2 called JSON script. The way we'd use this is we'd pass our dictionary without JSON encoding through to the template in the variable user here, and then we'd use the filter JSON script with some string as an identifier for the element that it should generate. Then our script becomes completely static.
Speaker 1: It could even be moved to a separate file And what this does is it retrieves the contents of this script tag using the DOM API getElementByID When this is rendered it looks like this, which can make it a little bit more easier to understand. So that first script tag coming from the JSON script filter, it declares its type is JSON. It has the ID that we passed it. And uh within that there's this uh JSON dictionary that we uh that that we pass through. But most importantly, the scripts, the HTML that defined the script tags, has been turned into
Speaker 1: Unicode element IDs in the JavaScript string. So these will no longer be parsed by the browser as HTML. They'll simply be seen later by the JavaScript parser and turned back into the corresponding elements. The code is using getElement IDYID to look up that script tag, look up the text content within it, and then parse that with JSON. I wrote a blog post on this because it's so important, so you could go check that out. As hackers, what extra can we learn here The first thing is to look through the generated pages on a site for inline script tags that seem to be passing data through. Anything that contains some data is potentially exploitable.
Speaker 1: And similarly to looking for other HTML injection attacks, we can try using a closing script tag at the start of HTML that we stuff into fields that we have control over. For protection, while we can use JSON script or other solutions that don't exist within Django And we can also use our friend content security policy again to block any resources that we don't trust. Going to squeeze in a quick bonus here for the talk after these stories. So if you are concerned about your site's security, there's a standard out
Speaker 1: that's fairly new called security. txt. The idea here is that you serve up a file, a well-known URL on your site with a machine readable um format for your contact information if there is a security vulnerability and this allows researchers who might just happen to be scanning all the websites that they're going to get access to to report issues that they find to you. To serve this out of Django, it doesn't take very much code. Here's an example view that serves the bare minimum security. txt that has a single line in it, that has an email address. I wrote a blog post on serving well-known URLs that uses security. txt as an example.
Speaker 1: You can go a bit more advanced and here's the for example here's the one on my own website which shows that it was signed with my PGP key and contains several different forms of contact. Nothing remains but for me to say thank you very much for listening to my talk. I'm Adam Johnson and these are my contact details. The slides for this talk will go up in this GitHub repository. Thank you.
Speaker 2: Yeah, I was uh thinking during uh I was thinking during your talk if something like you find another website is still in for example the Django admin um code in the HTML part or if someone have um scanned the HTML code in Django itself
Speaker 1: And so uh Tom Granger who I mentioned with the CAD emojis actually back in like 2015 or something, he moved to Django Admin to be fully content security policy compatible. So it's very unlikely there are any um major flaws in there anymore, like the thing with HTML injection when passing to JavaScript. Um there's definitely possibility like maybe some function in Django is using MarkSafe when it shouldn't. But I think these are generally it's it's pretty search-through. Like we have a lot of jet uh security fixes. Hopefully not.
Speaker 2: Okay, thank you. Ah
Speaker 1: hello Pascal Marius.
Speaker 3: Hey, great talk. Thank you very much.
Speaker 1: Thank you.
Speaker 4: Yes, the same for me. It was a great talk. Uh content security policies are really complicated. Do you have any advice or tips for ad isa small set of mas-havs that you should contain them
Speaker 1: Uh I think they're complicated because they like like an allow list mechanism. And you can get away by like you can still allow like inline scripts and like Still limit like the sources for your remote script tags, but that doesn't really add much protection. Um so yeah, it's kind of annoying to set it up Uh there is a Firefox extension that you can browse through a site and uh it will recommend a content security policy over out of everything that got loaded. Um so that can be a good first step. There's also it's report-only mode. So you can deploy a content security policy and have it report to a third-party server
Speaker 1: whenever a browser detects something was loaded that it shouldn't have been, and that can help you form. uh a good CSP. There are like a couple services that do this. I think Sentry does and um one by Scott Helm called uh
Speaker 3: Report.
Speaker 1: Report URI, yeah, exactly. Yeah. I don't know, Pascal might also have some experience setting this up.
Speaker 3: I'm just annoyed by it. I'm deeply disappointed. I I tried to do it for my website, but I it's still in the report only state because I included Google Maps and There's no point in activating it if you have uh Google Maps allowed because their JSON P endpoint just allows to for anyone to execute arbitrarily. JavaScript. And I'm really annoyed it was very complicated to get Google Maps running in the first place. And it's upsetting.
Speaker 1: Yes, it's surprising, isn't it?
Speaker 3: Yes. Come on, Google.
Speaker 4: Thanks. As you can see my kids are here so it's hard to ask questions.
Speaker 1: Thanks again. So
Speaker 2: Adam, you show the use of um JSON script tag. And the one I tried to suggest to Pascal before, but I think it's the main case the case he provided to to show to us it was a bit different Uh I don't know if there is some something not good in uh in line JavaScript. Yeah in the HTML page.
Speaker 1: If there's something not good.
Speaker 2: Yeah, because
Speaker 1: you mean in my example
Speaker 2: Yeah, but I suggest not to use uh inline JavaScript in the HTML. I I don't know if something in general
Speaker 1: That is good advice. If you're using content security policy, you'll want to activate its feature that disables inline scripts because they're normally uh emblematic of um an XSS attack um like this injection thing so like in my example I was just keeping it simple but yes it would have been better that I used a separate script file
Speaker 2: Okay, thank you. Now , I think I recommend that
Speaker 1: in my blog post, yeah. Uh Pascal, I was wondering if you have security d TXT files set up on your site.
Speaker 3: No, I did not. Actually I just learned from your talk about that. That's really cool. It's it's really cool, and that's definitely something I'm gonna share with my colleagues. Um that's definitely a takeaway And it's really cool. I will I will do that for sure.
Speaker 1: I think I've received one notification so far from someone scanning the web and they told me like I don't have a DMARC record set up for my email. I'll go do it I guess, but it was a bit spammy.
Speaker 3: But cool, yeah. I was very impressed when I saw that. I will I will for sure do that.
Speaker 2: There is something missing from the Django Project. com website
Speaker 1: What uh a security TXT.
Speaker 2: The security txt. Maybe we can add it.
Speaker 1: Yep. That that should be an issue on Django project record
Speaker 2: So I think uh your talk was the last one of the day.
Speaker 1: Indeed. Yeah, I think it's time to go join the closing remarks.
Speaker 2: Okay. Thanks for the same. So thanks again for the for the zoo.
Speaker 1: Thank you. Cheers.
An exposed debug page revealed enough information—including the session serializer, signed-cookie session engine, and secret key—to forge a session cookie. Because Pickle executes code while deserializing, the attacker could send a crafted cookie that ran commands on the server.
Discussed at 3:10Never deploy with `DEBUG = True`, run `manage.py check --deploy` during deployments, and avoid Pickle wherever possible because it has no genuinely safe mode for untrusted data.
Discussed at 8:35Different Unicode email addresses can uppercase to the same string. If password-reset lookup uses the normalized value but sends the reset email to the original input address, an attacker can register or use a colliding address and reset somebody else’s password.
Discussed at 12:29Upgrade Django to a version containing the fix and keep up with security releases. The vulnerability affected Django’s authentication code as well as GitHub’s Rails implementation.
Discussed at 15:40Marking a string safe after combining it with user-controlled names tells Django not to escape the input, so a user can store a script tag that executes when staff view the admin. Use `format_html` (or `format_html_join`) to combine trusted markup with escaped data.
Discussed at 17:59A Content Security Policy can block inline scripts and other untrusted resource sources, preventing injected scripts from running even if an HTML-injection bug exists. It can be developed using report-only mode and browser or reporting tools to discover required sources.
Discussed at 20:17An attacker-controlled value beginning with `</script>` can terminate the HTML script element before the browser’s JavaScript parser sees the data, allowing an injected script to run. Django’s `json_script` filter safely stores the JSON in a separate JSON script element and lets static JavaScript retrieve and parse it.
Discussed at 23:47`security.txt` is a machine-readable file at a well-known URL that tells security researchers how to report vulnerabilities. In Django, it can be served with a small view containing contact information, with optional details such as additional contacts and a PGP signature.
Discussed at 27:21Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025