Closing session
Published June 13, 2025
This video features Pradyun Gedam at DjangoCon Europe 2022 in Porto, Portugal.
Keynote: Growing pains of an open source project by Pradyun Gedam
A discussion of how open source software projects are managed, through the lens of an eventually-popular open source software project.
Pradyun Gedam explains how an open source project’s growth changes the work of maintainers and the structure of its community. As usage increases, issue trackers, support channels, moderation, documentation, communication, and contributor onboarding all require more deliberate processes; volunteer effort alone often cannot sustain them. Popularity also makes change harder: users create implicit interfaces and expectations, so projects need clear versioning, release communication, migration paths, opt-ins, and decision-making processes. He argues that healthy projects must invest not only in code but also in invisible community work, redistributor relationships, funding, and expertise such as user experience, because maintaining useful software at scale is fundamentally a communication and sustainability problem.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello, I'm Fradyun. I'm going to be talking about where Spain's that an open source project experiences. As it grows and once it has grown. Perhaps a more accurate title for this talk would have been a detailed description of the various factors that affect how an open source project and the community around it are maintained. But that doesn't really fit in one breadth, and it definitely doesn't fit on one side. A few quick words about me. As you've likely noticed, I work as a software engineer at Bloomberg, where I helped develop Software to make it easier for other software developers to author, develop, and deploy software written in Python. Outside of work, well
to use a metaphor, I wear a lot of hats in a lot of open source spaces. As is hopefully self-evident, I enjoy working within the Python ecosystem. With the introduction out of the way, let's take the first step in our journey of an eventually popular open source project. Creating the project and putting it up somewhere. These are surprisingly consequential steps Because they have a strong influence on how the project is perceived, how it is managed, and who finds it to interact with it. And chances are, for most people like myself, that isn't really a choice you're going to be thinking about today. You just use GitHub.
That's not a bad choice. I mean I pick it more often than not, but it's not the only choice. Not all of open source software development happens on GitHub And even when a project is hosted there, not all of the work happens on GitHub. Django itself is an excellent example of a project that's hosted on GitHub and is not using the GitHub issue driver Many others that aren't that well they aren't even hosted on GitHub. Projects can and do host their code on different platforms. And that matters when you're looking at the broader ecosystem of open source protects. I digress. Let's assume we've created a project repository on our favorite hosting platform.
Set up all the relevant automation around it and have done all the software work needed to create a pop project that's well ripe to become a popular one. I'm gonna skip over any discussion of what you need to do for that. Let's just assume my project has all of those pieces in place and all the stars are aligned. Let's also establish a definition of popularity. I'm going to pick a definition that a more popular project is one with more users. This means that nearly all projects written in Python are going to be less popular than Python itself, which makes logical sense.
Correspondingly, we can now compare two projects if we know the exact number of users they have. So phrases like more popular, less popular, they're meaningful. It's comparable and something we can measure. Well, we're going to be looking at indirect indicators of this to get a sense of how many users a project has. There's an entire topic around collecting usual metrics and doing it in a privacy-respecting way and how accurately they can reflect the actual user counts. I'm gonna duck out of that discussion. This is a good enough definition for our project, for our talk. And for a hypothetical project. It's going to become popular. It's going to become popular at some point.
That's sort of the whole premise of this talk. So let's look at how that will happen over time and what the implications of that would be. The growth curve could look like something that the Python package indexes had, uh IPI, which has grown four times in the last four years. This graph, by the way, comes from an excellent blog post by Dustin Ingram, one of the administrators of PyPI. To reiterate, this is an indirect metric. Whether PyPI becomes less popular over the weekend because there's fewer bytes being transferred over the weekend compared to weekdays is not a productive discussion This sort of exponential graph is what you'd expect to see when your project's user base is growing and they're all
interacting with the project regularly This is a project growing at a steady rate relative to its existing size. Exponential. A different curve might end up looking like the sudden spike in usage as the English version of Wikipedia had about 20 years ago, when they went from hundreds of edits to thousands of edits to millions of edits in the span of less than three, four years. This happens when a project fills a niche that wasn't being filled earlier, or through publicity pushes by being a project that's being published by someone popular or being publicized by some unpopular. There's a sharp rise in the various things the contributors and maintainers need to deal with, depending on which side of this curve they're on
Hopefully the software works just fine and the increased number of users isn't something that causes direct failures. However, everything else around the project also needs to scale as quickly as the user growth. That can be challenging, especially with a primarily volunteer group As we're going to discuss, the experience that the users, contributors, and maintainers of Wikipedia would have had is very different based on which side of that sharp birth they're on. I'll use Google Trends for Ruby and Ruby on Rails to well show that for certain projects, popularity is strongly correlated with another project or another piece of technology.
There's a certain way of using it. It's popular. And therefore it's popular. Another example, a hypothetical one, would be if Django adopted a new dependency in its next release When the next January release comes out, that dependency will almost overnight get a lot of new users. It will, for our definition, become a popular project. One way to become popular is to be used by popular things. One of the symptoms of increased usage of an open source project is that it's going to have more users Come to the issue tracker and other interaction points for the project. If the project's on GitHub or there's an established cultures
Users are going to go to the issue tracker and that is going to be the primary point of interaction with maintainers. Basically, there's going to be a lot more traffic that existing contributors, maintainers, will have to keep with, keep up with. Depending on how the project is set up, this alone may result in way more work than the maintainers and contributors can keep up with. For example, if you spent a 40-hour regular work week looking at the 900-ish issues on PIPS issue tracker You would have a little over two minutes to spend on average on each one.
That's not enough time to read the first comment in many of them. And many of these issues have hundreds of comments. And that's excluding pull requests that have both discussions and committee to review, which will definitely take more than two minutes on average. CPython has over 6,000 open issues and over 1500 open pull requests, merely keeping up with an issue tracker of a popular project. is clearly a non-negligible amount of work. I'm aware I've picked somewhat extreme examples in terms of popularity within their ecosystems. And I'm literally mentioning CPython, which is the software you download from python. org
Partly because these are projects I'm the most familiar with. And partly since they help make the point I'm trying to make. Keeping up with the issue tracker is a lot of work. Which is why it was a dropping point for the internet a few months ago when Flask managed to hit zero open issues and pull requests. That's an achievement, all right? Not only did the maintainers actively keep up with incoming requests, they also worked through the backlog. That's something that took many, many hours of effort, carefully fair responses, and as I'll discuss later as well, a willingness to say no. One of the best ways to prevent your issue tracker
from being flooded is to answer the user's questions in other places, like dedicated project documentation or a different dedicated photo. For documentation, well, we're in the internet age now. So this ought to be a website that's publicly accessible and indexed by search engines As it turns out, this is how the vast majority of a project's users will interact with the project. Thankfully, ReadThe Docs, Netlify, Versal, many other static website hosts. make it relatively straightforward for most open source projects to just publish the documentation as a static website. Of course, I'd be remiss if I didn't mention this. Django has one of the best documentation sites in the Python ecosystem.
And I think it's kind of neat that it's a Django application itself. Documentation is where you should be answering the most common questions about the protect. And hopefully you will no longer have to answer those in the issue tracker. At the very least, you'll have a link to point to instead of having to repeat yourself. Depending on the size of the project, when it rose to popularity, who the audience is, it's quite possible that the issue tracker is not the only place where people are asking questions So having the details of how things work and the common questions, having these details. in the documentation
means someone else can point the user to it or the user can search for it and find it themselves, reducing the load on the contributors. And I do feel like I should mention that access. I hope Danielle is smiling seeing this slide. I find DAND Access is a useful framework for how to think about documentation when authoring and improving it. That documentation is not just one thing, but for distinct things based on what the user is trying to do. And understanding the implications of this, how each of those four different things work, can help improve most project documentation. That's almost a quote. I'm not going to go into the details of the framework and how to apply it.
That I'll point you to the very good looking website that's linked in the upper right-hand corner. Where Daniel's done an excellent job of explaining things. I'm pretty sure I could fill the rest of my time just talking about documentation stuff and how this connects to various other topics like communication, discoverability. onboarding contributors and so on. But not today. There's other stuff I want to cover. I do want to jump back to C Python though, which is about as big a project can get in the Python ecosystem. To talk about a vital piece. Certain kinds of discussions are not conducive to the issue tracker format, not just because of the potential for overload that we just discussed
But also because the format just isn't the right one. For Python, there are a lot more communication channels that are not CPython issue tracker There's the discourse forum, discourse. python. org. So many mailing lists on mail. python. org. Mailman. And there's a public discard server, a Zulu instance, and more. Even Python enhancement proposals are a medium for communication. One that's effective enough to have inspired similar processes across the Python ecosystem, Django's DEPs, NumPy's NEPs, as well as Well, many other language ecosystems processes as well.
And I've only mentioned channels that are directly maintained and managed by the project's primarily volunteer contribution. contributors and these channels represent a lot of upkeep effort. There are other channels, podcasts, blogs, conferences, and so much more. And in addition to the monetary costs of setting these up and the infrastructural costs of well keeping forums online and mailing lists online. There's also a non-negligible amount of work involved in ensuring discussions there happen smoothly And those discussions, those interactions play a large part in who participates in the project.
This is an excellent infographic created by Alex Bailey about how a person moves. From having never heard of something to being deeply involved in that thing. I really like this infographic. Let's run through it, through well, my personal story. To just illustrate how it works. I started learning Python in 2013-ish when it was already a somewhat popular language. I gained awareness of what eat and paper well through a book that I was reading to learn the language. I used these tools a bunch and got some understanding of them. After some time I ended up on the issue tracker of the project and reading things
eventually posting uh a summary um the things I read on that specific issue in that issue tracker. And well it was a well-received one Everyone was aggominating of the fact in the discussion that ensued after that I'm a teenager who is engaging in a technical topic I don't fully understand. Uh I didn't have the ability to stick around though, having grown up in a family that has access to computers, internet, good education, and sufficient means that I could just spend time on the computer, on the internet. And the more I worked on the project and collaborated with the maintainers on various issues, the more I felt like I could, you know, stick around. Doing our Google Summary of Code, GSOC, in 2017
helped a lot with that. And a few weeks after That project, I excitedly said yes when my mentor, Donald Stuff, offered me the opportunity to become a maintainer on the project. Get the commit bit One thing I'll note is that the number of people in each of these stages going from top to bottom, it reduces like a funnel. uh the difference between the number of people who are aware of a project and the number of people who understand what it is to the people who have identify themselves as being able to work on it. It's going to be an order of magnitude and more.
In other words, there's going to be a lot more people in the earlier stages, the higher stages of this pathway than the later ones. Some may not even cross the stage, may not be able to cross the stage without help or assistance. I was lucky enough to get assistance each step of the way through my pathway to ownership of well In a similar spirit, a project is going to have a lot more users who read comments on forums, issue trackers, TACOVE flow. than the number of users who will ask those questions or respond to them.
Moderation, spam, code of conduct issues. These are not things people get excited about. I certainly don't. But these tasks are vital to undertake in order to ensure that you have a healthy community. These play a big role in setting the tone and precedence for what is acceptable and how you can create spaces where we're able to help our karate and future friends within the community. It doesn't really matter what the size of the project's user base is. The behaviors that are deemed acceptable within the project's communication channels will determine who and how many will want to continue to be in those spaces.
Another thing that is affected by increasing popularity Is the ability of a project to make changes. Well, certain kinds of changes. I think it's a safe bet to say that many folks in the audience have had to make changes to existing code. when it's either being published and deployed or is being actively used. Depending on what the expectations are, you need to make changes in a certain way to match those expectations. If it's okay to have the application go offline for a few hours every weekend, you can take approaches that you wouldn't be able to take if the application has to be online for 99. 999% of the time. There are some battles to that with open source projects.
How you make changes is determined by what expectations are. uh what the expectations are and what approaches do you have established for making those changes More often than not though, you will need to make changes. There are very few pieces of software that don't need to be updated or improved in any way. That brings me nicely to the stable interface paradox, an excellent idea from a keynote by Paul Gansel, whose work and ideas have already been mentioned in this conference. A project that's going to be a project is going to get a lot more feedback about it after it has a large user base. But that's also when the ability of that project to make changes
without being disruptive decreases dramatically. In the early stages of the project, when it has few people developing it, few people using it, there's you can make drastic changes to it. But you also don't get a lot of feedback about the design choices made and whether the name you've picked is a good one. You'll know If you got these foundational pieces right, as you get more users, and even if you did get these pieces right, You'll still find new bugs and some bugs can't be fixed in a backwards incompatible way. We saw this recently with the most recent Python releases where there's a new backwards incompatible security fix
When you make a change, you will break people's workflow. At a certain level of popularity, it's no longer a case of if, it's a case of how many And it would be wrong of me to not include the relevant XK CD panel about this while talking about breaking changes. There's actually a law for this, Hiram 's law. I don't know if I've said names right this entire slide. I apologize. And the reason I want to point this law out is because I've been bitten by this many times. And chances are decent
that there are people in the audience who have noticed the instances when I've been bitten by this because they were depending on the observable behavior that changed. This can become a big source of friction and frustration for both the contributors to a project and well the users. For the contributors and maintainers of the project, they have to deal with the vastly larger surface area for what the users expect to be using. uh what they can expect the users to be using and what they can safely change is a shrinking group While the end users have to deal with unannounced and unmanaged, as we'll discuss,
unmanaged changes that may be disruptive to their use. Basically, you can't know all the ways users are using your software. This becomes more and more of a problem as a project grows. You can estimate it. You can Make an educated guess. You can even conduct user interviews by and potentially have experts do those interviews. But you're still going to have A sample of users. Still not gonna cover everyone. This limited visibility, not being able to see the exact way your users are using the code. Is a major contributor to the problem of the out-of-contract use
and implicit interfaces that Aramslaw talks about. All this is to say, as your project becomes more popular, you'll have to invest more and more time and energy into change management. I like to think of this as a mix of communication, setting expectations, and providing support. Version numbers are part of the communication here. There's some work which prescribe exact semantics to the components of the version and it sets expectations for what those mean. This also makes the most basic approach, which is just bump the major version. That approach can work really well.
It sets the expectation that there's potentially disruptive changes. and also communicates that the project intentionally made that change. There's a lot more you can do though. You can invest in effort into writing clear documentation. uh giving conference talks, discussing what's changing on social media, podcasts, uh publishing release announcements about changes you've made Exactly. There's a lot. And a lot of this will be picked up by people who don't necessarily have the commit bit. These are contributors, not good contributors.
A goal with all of this though is to communicate to the users there's going to be a change Provide resources to support them and to clearly set expectations for what that change is going to look like. To what extent you need to do this depends on Well your willingness to do it at some level and also what is deemed acceptable by both the maintainers and the users On the other end of the spectrum, you can design your processes to embrace change and communicate this. Set up processes and expectations based on that Out of parallel, uh Django 's migrations are a great way to embrace that
your database structure is going to change. It establishes a process and manages expectations around how you make those changes If you could make these manually and have the same effects, but this framework gives you something to think about. This gives you a clearer model for thinking about this and for managing this And it's useful. I actually really like Rust 's approach for handling changes to the language. They have a nightly build which has a bunch of early access. features or I guess changes to the language that users can install and manage using well supported mechanisms. And typically
these fun features, the nightly only features, are additions on top of existing stable functionality. The new things being added. And they don't affect code that already existed. They also don't make any backwards compatibility promises about these nightly only features. That's the expectation. They will change. Allowing for those features to evolve based on feedback that enthusiast users, eager users are providing by using this nightly release. This feedback cycle eventually will end and the feature will stabilize, be enabled by default in the stable role to be a part of it
Where they'll now be subject to a stricter backwards compatibility promise. And the stable releases happen often, every six weeks. There was one yesterday. Rust also has additions, which serve as their points for making backwards incompatible changes, like adding new keywords And they allow for code written in different editions to be used within the same program, making the migration between editions softer. Now granted, the amount of effort is going to be needed to that you need to put into the change management work for programming language is going to be greater than
most projects that are not programming languages. The thing I really like about this approach, beyond the fact that it embraces there's going to be changes and manages them and sets a clear expectations It makes it possible to add new functionality on a regular basis, making it available to everyone with strict compatibility promises. It gives contributors a short feedback loop from enthusiast users who will try out new functionality. And it provides strong compatibility promises for older code and and users who care about that Churn is
a marketing term for basically users who stop using something. Depending on what the project does, it has a churn budget associated with it, how much change users are willing to deal with to be able to use that piece of software and stay up to date with it In other words, how chaotic a project can be before users walk away. A certain class of users won't care about most of the changes And then again, given sufficient popularity, there will be users who want that change to never happen or to have never happened. This is complicated because this makes it, well, again, difficult for the project to evolve and move forward
to make changes. Any sufficiently popular open source project will need to have some sort of framework for making changes. And it'll need to be mindful of how it utilizes the churn budget. Versioning handles a lot of this, especially the communication around uh well changes in churn. Opt-in, opt-out mechanisms can help make cautious use of this churn button, allowing you to make changes and then tell users to flip flags to get different behaviors. The decision between making it a configuration variable versus a temporary opt-in, opt-out depends on what the long-term plans for the project are.
If it's something that's going to be added and then become the default, that's going to take a different approach than if it's something that's going to be removed and the users need to migrate. But they can use it for a little longer, you know. Another thing that will push you is deferring user needs. What is essential for one user can be problematic for another. And as your project becomes more popular , there's going to be more of this. There will be users who favor stability over new features While others who really, really want this one more new feature. There's going to be multiple forces pulling in multiple directions. There's going to be users who are happy with the current state of affairs.
But they will be noisy if you change things. But there's a currently noisy group telling you that they want to change. There will be I need to start bundling users into groups based on what they need. Uh which for maintainers to sort of keep a track of this user personals right and well this is also where learning to say no is important Contributors are going to be operating with limited context on what the right choices are. And this makes the decision-making process fuzzy.
Once we start making decisions as a group though, we need to have some sort of social structure for making those decisions. Early on, leaving this implicit is fine. And many projects operate on the project creator is makes the final call. uh the benevolent dictator for life, the BDFL model. It it works. It worked for Python for a really long time. In software age terms anyway. And As more and more decisions are being made, as the number of contributors grows, there will come a time when
the process needs to be written down. This is a vast and complicated topic. What works for each project and each community is going to be different. It will change over time. Some will want to operate on consensus. Others will give special privileges to certain individuals. Yet others will, well, what? Having this process written down in a meaningful and useful manner does make things easier for everyone. But the process of writing it down for the first time and even changing it later won't be
I won't describe it as fun. I will describe it as worthwhile though. It makes things clearer for everyone and perhaps importantly Uh it reduces uncertainty, which is valuable when you're onboarding new contributors to the project. Invisible work. Honestly, Kojo covered a lot of what I wanted to say about this in his opening keynote for this conference. And I will just point you to his talk. And well, if you're watching this online, go watch Kojo 's keynote after this one. On the other hand, I get to cover other topics with a little bit more time.
As your project grows in popularity, there are going to be redistributors who repackage it for their ecosystem, let's call it, and ship it to users. They play an important role in the broader software ecosystem as well as consumers of open source. And there are a lot of redistributors. Here's a snippet from a pull request discussing how a piece of software is repackaged by many redistributors. That's patchling, by the way, of PyPI project Redistributors play an important role in making a project accessible to audiences that wouldn't be able to use it otherwise or
can't use it in that specific context otherwise. They are the reason that Python ships on nearly every modern computer out of the box. And why you can type Python on the terminal and have a Python interpreter. Almost regardless of what flavor of operating system you're on. They're also supposed to serve as a layer between the Project and X users handling support requests, for example, from the end users This lets them cater to differing user needs. Red Hatch Enterprise Linux serves enterprise users with enterprise supposed contracts. uh
or different user personas which like SPAC which started by serving the needs of large supercomputers uh the large supercomputing use case the HPC use case In reality, things are a bit murkier, maybe a lot murkier. For one, the goals of a redistributor may not align with the goals of a project The policies may not align and the redistributor may modify your code so that it behaves differently. uh breaking assumption somewhere. Hiram 's law. And well they're on the hook for that, right? They they're supposed to be an intermediary
between the project and the end users and in theory they would handle this. But sometimes it's easier for users to reach the project itself rather than the redistributor who they're getting it from. An example I can give for this that I'm guessing nearly everyone in the audience has hit at some point, or at least the majority have, is how How Debian modifies Python. Their modifications to Python actively make it painful to use it as a software developer. And that's because Debian doesn't want the Python that they distribute. It's not meant for end users, except no one's telling the end
users that and hey you can get it as part of your operating system is well part of how Python is popularized. And worse still, things aren't just plain broken. They're broken in subtle ways It's not that the commands you run don't succeed. It's that sometimes they do the wrong thing without telling you. It's a degraded experience And changing this is a complicated people problem rather than a technical one. It's a flash of what different users need or different groups need and how changes are being made. Those are complicated problems.
Sustainability on its own is worth Extended discussions and I'm not an expert on this topic. So I'm gonna keep this brief and not go into an extended discussion. One bit I feel like I should mention is Funding programming tasks isn't enough. We need to fund the work that happens outside of writing code as well. I hope this you know establish some of why. Uh it funding this external Funding new contributors or existing experts who are currently external to the project and have skill sets that existing contributors don't is extremely valuable. This was something I learned firsthand while working on PIP as a part of
Grand Funded team in 2020, where we had user experience experts. working alongside programmers uh on various aspects of PIP , primarily the resolver And I hope it's self-evident , if you have a popular project, popular open source project, it takes a lot of effort to keep it and the community around it Functional and healthy. The bulk of this work, at least today, is being done by volunteers in their free time, which as I hope is self-evident isn't sustainable. And the foundations like the DSF, uh
Django, uh Software Foundation, and the PSF, the Python Software Foundation. and many more have their own fundraising efforts and effort towards sustainability for the projects that they support. Where, well, people who have spent more time in open source spaces than have been alive are thinking about these problems and actioning on them There's lots of reasons to be engaged in open source. And at the end of the day, for nearly every open source project, the code we write is meant to be useful to people. The people it's useful to, it could be just yourself, your immediate peer group, or complete strangers on the other side of the world And as the number of people this code is useful to grows, the project will need to adapt and grow as well.
To code carabilling, great code requires communication. There as there are more people, you will need to communicate with more people to continue to build these things, to make the software and the core better. to to make the community you find yourself in better. And well to me that's what makes open source great. It's that you get to work with many awesome people to build cool things. Thank you. List of references at the end of the slide deck.
More users generate more traffic in issue trackers, forums, and pull requests, creating work that can quickly exceed what volunteer maintainers can handle. Even merely keeping up with a large issue tracker can become a substantial ongoing job.
Discussed at 7:09Documentation should answer common user questions in a publicly accessible, searchable form, so users can find answers themselves or contributors can link to them instead of repeating explanations. It also helps people understand how the project works and reduces unnecessary issue-tracker traffic.
Discussed at 9:29Issue trackers are not suitable for every kind of discussion, especially as a project grows. Forums, mailing lists, chat, proposals, blogs, podcasts, and conferences provide better formats for different conversations, though maintaining them requires significant effort.
Discussed at 11:38Moderation, spam handling, and code-of-conduct work establish what behavior is acceptable and help keep the community healthy. Those norms strongly influence who feels able and willing to continue participating.
Discussed at 17:23A project receives much more feedback once it has many users, but at that point changing its design or interfaces becomes far more disruptive. Early in a project’s life, maintainers can make drastic changes easily, but they have less feedback about whether those choices are sound.
Discussed at 19:49It needs deliberate change management: communicate what is changing, set expectations with versioning and release announcements, provide migration and support resources, and use opt-in or opt-out mechanisms where appropriate. Processes such as Django’s migrations or Rust’s nightly releases can make change more predictable and provide safer feedback cycles.
Discussed at 22:55As the number of contributors and decisions grows, informal arrangements such as a single project creator making the final call become insufficient. Writing down how decisions are made reduces uncertainty, clarifies expectations, and makes onboarding contributors easier, even though creating or changing the process can be difficult.
Discussed at 31:21Redistributors make software available in ecosystems and contexts the original project could not reach, but their goals and policies may differ from the project’s. They may modify the software in ways that break user assumptions, while users may still seek support directly from the upstream project.
Discussed at 33:44Funding only programming is not enough: projects also need support for documentation, community work, moderation, user experience, onboarding, and other non-coding tasks. Much of this work is still done by volunteers in their free time, which is not a sustainable model as projects grow.
Discussed at 37:41Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025