The promised Django Land; the tale of one team’s epic journey... by Nicole Zuckerman
Published October 25, 2019
This video features Nicole Zuckerman at DjangoCon US 2018 in San Diego, California, USA.
DjangoCon US 2018 - One Engineer, an API, and an MVP: Or how I spent one hour improving hiring data at my company. by Nicole Zuckerman
The fact that tech is struggling to hire or retain employees from diverse backgrounds has been written about and discussed thoroughly, particularly in the last few years. The economic, societal, and moral benefits of diversity are also well documented. Why is it hard, then, for well-intentioned organizations to shift their demographics? There are a number of reasons, but one that doesn’t appear to have been thoroughly discussed already is the challenge of gathering and responding to data about diversity within a company’s hiring pool and existing employees. One hack day, I was involved in too many projects and had only a token amount of time to devote to the one I was most interested in; seeing if we could determine whether we had sufficient diversity for any given role to start interviewing candidates, or if we needed to spend more efforts sourcing diverse candidates for the pool. I accomplished an MVP in approximately an hour, once I had an api key and permissions. It doesn’t necessarily require a huge effort to make a big difference.
In this talk, I’ll walk through the MVP indicating, for each role, whether there was “sufficient diversity”. I’ll also address gotchas, limitations, and What Now.
This talk was presented at: https://2018.djangocon.us/talk/one-engineer-an-api-and-an-mvp-or-how-i/
LINKS:
Follow Nicole Zuckerman 👇
On Twitter: https://twitter.com/zuckerpunch
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Nicole Zuckerman describes building a one-hour Django MVP that pulled recruiting data from Greenhouse, aggregated applicants’ voluntary EEOC information, and showed diversity statistics for each open role. She explains how a production version should avoid slow live API calls, restrict access, anonymize data, suppress small groups and combinations that could identify people, and exclude active candidates so the data cannot influence illegal hiring decisions. The data is incomplete and difficult to interpret, but it can reveal where underrepresented candidates are missing or dropping out; she argues for aiming toward labor-force parity, involving diversity, recruiting, and legal experts, and using the results to improve sourcing and each stage of the hiring pipeline.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Uh hi, I'm Nicole Zuckerman. That's me on the left. Best wedding picture ever, right? I'm a senior software engineer at Clover Health. Yes, we are absolutely hiring. Let's talk later. I've been interested in issues of representation and diversity for around 16 years now, and I love that my technical skills allow me to affect how we go about doing things at Clover Health and improve diversity inclusion. It's like infinite cosmic power. So I'm glad to be here talking about it with y'all today. And if you want, you can like message me on Twitter. That's my Twitter handle. That's right. I never post though, so I do check it. Anyway, uh
we know diversity is valuable. What stops us from collecting data to help us improve? The fact that tech has struggled to hire and retain employees from diverse backgrounds has been written about and discussed extensively, particularly in the last few years, and the benefits are also really well documented. So why is it hard for well-intentioned organizations to shift their demographics? There are a number of reasons, but one I don't think has been discussed as thoroughly is the challenge of actually gathering and responding to data about diversity within a company's hiring pool and existing employees. If you want to talk about getting that information from employees, let's talk later. For today, I'm just going to focus on getting that data in your diversity, in your Candidate pipeline. I know words. I have a master's in English literature. Of course I know how to speak.
So one hack day, I was involved in too many projects and had only like a token amount of time to devote to the one that I was most interested in, which is seeing if we could determine whether we had sufficient diversity for any given role to start interviewing candidates, or if we needed to spend more time sourcing diverse candidates for the pool. I accomplished an MVP in approximately an hour once I had the API key and permissions. Honestly, getting the right permissions from HR was like the hardest part, sorry, from recruiting. That took the longest amount of time So if there's anything you take away from today, aside from picture of my dog, which I'll show you at the end, I want it to be that it doesn't actually take a huge effort to make a big difference and everyone should give it a go.
Alright, so here's what you're in for. I'm gonna walk through the MVP that I did during hack day. I'll talk a bit about the V2 I'll talk about some gotchas and limitations that I uncovered, and then we'll talk about what happens now. You in? Alright, let's do it. Alright, so uh this is the web app I put together for Hack Day. It was around a year ago. This is the MVP. The front end literally has no styles, is as bare bones as bare bones as can be, but it works I am not a front-end developer by any stretch, and you'll be able to see that in a couple slides. Alright, so there's uh it's a simple Django web app with one view. There's a bunch of real-time API calls to greenhouse
to fetch data. Greenhouse just happens to be the um recruiting funnel tool that we use. There are other ones out there that also have APIs, but this This just happens to be the one we were using. And this is pretty hacky code. In production, this would never fly with like unnested for loop. And like the performance of the code is terrible. But as a proof of concept, it was really effective to give people a chance to see exactly how homogeneous or diverse our candidate pool really was for any given role. So this is the one Django view. At the high level, it gets a list of jobs, it gets all the applications for each job. And then for each application, it gets the EEOC data associated with each application.
Whether you say that you're a man or a woman, what your race is, whether you're a veteran, whether you have a disability. It does a little number smushing at the end and renders a list with of jobs with data about how diverse the candidate pool is for each. So this is the thing at a high level. Here are the API calls I make. They're not fancy. I hard-coded the URLs because hack day. I hit greenhouse's API and load the data for each of these three endpoints. There's one where I get the list of jobs that are open right now and were created after a certain point, created after, because um When I load this in a web browser, it literally takes 10 minutes. Hack day.
And then I for each of those jobs I get the applications. And then there should be one or zero EEOC uh chunks for each application. And so here's where I add an applicant status to the blob for this job. I anonymize and tally number of applicants self-attesting as belonging to a protected group, and then just display aggregates for each role on one page. So this will say there are blank women in the pool for this role, or there are blank veterans. Front-end
developer, I am not. So this is what the thing looks like for any role. It says whether there is sufficient diversity, we'll talk about that later, to move forward with interviewing candidates for the role. Alright, so now that we've gone through the hack day version, what happens then? There were a lot of things I probably needed to change. to make this uh ready for production. The one that was killing me was those nested for loops with like real-time API calls that took 10 minutes. So this is like me hitting the endpoint, the um the view once and then there's just the um output goes on for days and that's just like hitting it one time
And you also don't want to have people waiting 10 minutes to see their results. So because of the way that uh greenhouse's API works, I couldn't do it together any more than I already did. But this is a ton of API calls. It'd be much better to collect the data regularly, store it in a database where we can make fewer queries by using table relationships. So our first order of business was to sync the new data from greenhouse to our database and have no synchronous API calls to a third party. and then our page can load in a snap. Now the next hurdle is how do we know what data is new? How do we upsert instead of getting everything fresh each load? But that's a problem for future, Nicole
Another those are not mine, I wish. Another thing we decided to do for V2 was to only show diversity stats for candidates who have exited the pipeline already. people who have either accepted an offer or left the pipeline because they withdrew or we decided not to move forward with their candidacy Otherwise, there's potential for a hiring manager to look at this page, see the candidates we're evaluating, and go, oh, we need more veterans to increase our EEOC data for this year. We should totally hire this woman because it's illegal. You can't make hiring decisions based on someone's protected class. Like whether you know being female or veteran status, disability status, and race.
We also added a list view of all roles where you can click through for details for a particular role or department. The other things that we wanted to accomplish with V2 were showing diversity at every step of the pipeline from recruiter screen through offer acceptance. And being able to slice data to see the story for a particular group or from a particular source or from a particular recruiter. So you can say like, oh, we get really great candidates from this source, but this source not so much. Or we're really we're super represented with people of color, but we're not really great at reaching out to people with disabilities Alright, so here were the things that totally tripped me up.
Fortunately I figured stuff out or had people help me figure stuff out before any sensitive data anywhere. Any sensitive data went anywhere. So one. Be careful with people's sensitive information. I suggest you host the data where not everyone can access it, like this poor puppy. But you know that's it's it's probably made of chocolate and he shouldn't have it. Um so one thing that we did was we anonymized our data when pulling information about candidates using an API, for example, we didn't store any information that's personally identifiable. No names, no nothing. I also recommend that your database not be in a place where all of engineering can access it because these are like potentially their
coworkers or themselves and it's not really cool to surface People's personal data for other people to access. So I would also say put your API keys in a private place that not even all of engineering can access. Otherwise, if you have your database stored in a place that is private But everyone can access the API key, they can get the data for themselves, and then where are you? And I would also argue for different levels of permission, be sensitive about who can see what. Maybe everyone should get the 10,000-foot view in an all hands, that's the cookie version. But Being able to see per role slicing on different tags should maybe be limited access and maybe only like your diversity and inclusion manager, if you have one, gets to see like all the granular stuff
That's the ingredients. Metaphor. It's really only a meta-three, let's be real. So another decision that we ended up making was to show results only for groups of minimum pool size or greater, even though that means that we get data on a slower cadence than real time, in order to give the pool size a chance to get large enough for real anonymity. Be aware of combinations. Veteran, disability, gender, and race combinations are identified identifiable even with anonymity. So for example, if I'm like, hey, who were the musician Muppets? There's like five of them. And if I say, who are the blue Muppets? Yeah, there's like five of them. But if I say, which are the blue musicians?
There's only one. He's very easily identifiable, and you don't want to have that. So let me just stress for you one more time the importance of um anonymizing data and not after you um keep your minimum pool size And also consider not including candidates who are currently in the pipeline. Make sure that No one can make decisions based on any of the about a candidacy based on the information that is in your web app. Like I said before, that's super illegal.
Right, so I was telling you not to do anything illegal, right? Yeah, that's totally it. Um and making any decisions about a candidate, whether you move forward or not based on their protected class is totally illegal. Do not do that. And so the best way that I found to do to not do that was to not give was to not give people the opportunity to do that and show information for people who are currently viable candidates. Cool? Puppy in a pipeline. Okay, so another challenge that we have is that um not is uh data sparsity. Not everyone fills out information. It's not available necessarily for referrals or candidates who are sourced internally. And even on the website, not everyone chooses to fill out this information Even though in my head I was like, yeah, I'm gonna get this data and it's gonna be glorious.
We're gonna have so much information to make decisions on, and then you get to reality and like one person has filled it out So under um understanding why the data is sparse will help you figure out how to mitigate it. And understanding why the data is sparse has a lot to do with humans. Solutions regarding humans are usually specific to a group of like what they have in common. So like what works for employees is not the same as what works to get more data from candidates, for example So sometimes there's like no opportunity collect to collect data from non-system inputs, for example, the source candidates. So what you can do about it is instrument human systems like have your recruiting folks make sure that they tell, that they ask for this information with all candidates
Also, folks are often suspicious about what you're going to do with their information and are worried that it's going to be used against them in some way. Since you don't have a rapport built with them very much by the time they're applying on the website, there's not necessarily a ton you can do about that except maybe like put a line saying, we respect your privacy and your data. This information is going to be used for like these purposes. etc just to like give them a little confidence that you are not a jerk, right? And you're not. So then another problem is that the data that we need to collect by law doesn't necessarily match up with what we want to collect or what people might be willing to divulge. We're also we're just limited by what it's legal to ask. Like you can't ask anything about someone's age. And you also shouldn't ask anything that's gonna like
lead to you figuring out someone's age, because um intent matters. Um And you also need to decide what labels and groups to include, like for gender questions, what are your options? The federal government collects only male or female or let's think men and women or something like that, but like I know that there's non-binary people out there. How can I reconcile the legal requirement from my understanding of the world and wanting to be inclusive? Making a survey is an art unto itself and it's possible to ask questions and elicit a response that you did not intend. So if you have like researchers who work at your company Maybe ask them for their help.
Federal contractors have an EEOC requirement that's Equal Employment Opportunity Commission that requires these certain questions and potential answers that don't certainly don't line up with what I would like to ask. And so that can lead you to have like multiple surveys, which causes survey fatigue and again suspicion. Alright, so now let's get to this elephant in the room. What is sufficient diversity anyway? There was this article in the Harvard Business Review about um the relationship between what the makeup of your finalist pool for a role and the hiring decision. So if you have zero diverse candidates in a pool, you're obviously not going to hire anyone from an underrepresented group.
But according to the handful of studies that were in this Harvard Business Review article, even with one candidate from a minority group, that person has like zero chance of getting hired because of our preference for the status quo As the article says, deviating from the norm can be risky for decision makers, as people tend to ostracize those who are different from the group. For women on minorities, having your differences made salient can also lead to inferences of incompetence. Basically, being the only woman in a pool of finalists highlights how different you are from the norm and then makes you riskier. So this also works for racial minorities. The odds of hiring a minority were 193 times greater if there were at least two
minority candidates in the finalist pool controlling for like the number of other minorities and white finalists. And the effect held no matter what the pool size was, like six finalists, eight finalists Um so like even having two people from an unrepresented group makes a big difference. But that's a pretty low bar, right? What we saw was from like a limited size of finalists. So if you've got an applicant pool for a role like software engineer, you're gonna have way more candidates to choose from. So even if you have two people from a minority group represented, if there are a thousand people from the in-group Even putting aside unconscious bias, which totally exists and will whittle the diversity of your pool at every stage of the pipeline, and our status quo fix of having two candidates, you're statistically still very unlikely to hire someone who's from the status quo, who's not from the status quo.
All right, so what do we do now? Why not try going for parity with the eligible label labor force in the US? Then at least if the top of your funnel is balanced, you stand a greater chance of having those people make it through your hiring process Not that it's smooth sailing from there. We still need to institute systems to limit the effective unconscious bias and the rest of the hiring pipeline, but it's a start. So this is a table of gender and race in the potential labor force from 2017. The base number is a thousand, so the total number of black women who were in the workforce was 10. 5 million. Out of 17. 5 million black women who could have been in the labor force, approximately, I round , not all those people are qualified for every job or even interested. But if we support programs that give training and opportunities to underrepresented groups, while we also try to get as close as possible to having our pipelines represent that population pool, we might actually get there.
But if we don't set our sites that high, we're never going to All right, so what now? Well it turns out we found out when we were midway through our um V2 that there's a company that does a lot of the hiring uh pool stuff. f out of the box. So we ultimately decided to use them for the future. So engineering hours don't need to support the web app indefinitely. Could we have done our research to discover this company first and save ourselves the work? Absolutely. Would anyone have thought to Google it before I made my MVP? Unlikely. We found out about the company through our diversity inclusion manager who joined right after I made my MVP. Of course. If you have D<unk>I people, talk to them early. If you don't have the budget to outsource the way we ultimately did, the web app
we discussed is perfectly usable. It probably requires some further iteration to handle the sparsity of data, and you need to figure out what is sufficiently diverse for you. But even if your time is limited, the MVP is better than nothing at all. If you go this route, make sure to include any D<unk>I people at your company and recruiting and legal. But if you're willing to do the work, most of the time people are happy to give you their opinion and send you off on your way. But don't try to do this without knowing what's legal. Someday it would be really great to integrate data from candidates with data from current employees. But that's pretty challenging because those are almost never in the same systems. So maybe next hack day.
Now when you identify different departments or roles where there isn't representation across different groups you can go back and source those missing groups. If there's a big drop-off at one step of the recruiting funnel, investigate why people are leaving there. So this is fake data, but this is a format of a report I look at quarterly. For one role, software engineer, you see the counts and pass-through rate for men versus women for each stage of a hiring pipeline. At the beginning you start with 100 men and 50 women in the pipeline, which is not representative of the population, but not awful. But in the application or resume review stage, 50% of the men make it through, but only 20% of the women. Looks like you need to see what's going on in that resume review stage. At the phone green stage, you still have more men than women, 50 and 10 respectively, but they make it through at the same rate. 30%, that seems okay.
Then for the on-site, you have 80% of men making it through the stage, but only 33% of women. Uh-oh, maybe take a look at your on-site and see what's happening there. And because I work in a really great place, everyone accepts their offer, but wow, we lost a chance to hire women at the resume review and on-site stages So this is hard. Do it anyway. Since the problem is one of people and systems, it's inherently challenging, especially if you're a person who's at the apex of privilege in our society. Do what you can, from my little perfectionist heart to yours. Admit you won't have it perfect. Try to build processes around data collection wherever you can. Don't guess about someone's gender or race unless you're required to by law. Or if you're going to, which I don't prefer that you do, outsource it to a company that does it professionally like with machine learning.
They're honestly like Probably slightly less biased than humans. And if you are at the apex of privilege, cishet white men You should involve people from underrepresented groups to be involved in the process. Listen to them. They know how things come across. But don't make them do the line share of the work. My recommendation is to solicit their opinion early and often. Do some work, show them what you did and ask for feedback, but also deal with it with grace if they don't want to be involved because they've been carrying this burden the whole time up till now and you're like coming to the last mile, like come on So those are just some examples I happen to have. There's a ton of other things you can do to instrument your hiring pipeline or measure changes in inclusion at your company. Share out what worked, what didn't work for you so we as a community can learn from each other and develop best practices.
There's no reason to keep best practices about diversity and inclusion secret. I want the special sauce all over tech So that's me and my Twitter handle. That's my dog Chloe. If you uh follow her on Instagram at Corgi in the Front, I promise you will not be disappointed. Her front legs are a little shorter than her back legs. And I work at Clover Health. It's really great. Come work with me. Chloe says you should work with me. Thanks.
Nicole built a simple Django app that queried Greenhouse for open jobs, applications, and EEOC data, then anonymized and tallied applicants who self-identified with protected groups for each role. She built the proof of concept in about an hour once she had the API access and permissions.
Discussed at 3:19Instead of making nested, real-time third-party API calls on every page load, regularly sync the recruiting data into a database and query the local data. The V2 also showed role details and limited statistics to candidates who had exited the pipeline, reducing the risk of hiring decisions based on protected characteristics.
Discussed at 6:22Anonymize data and avoid storing personally identifiable information, restrict database and API-key access, and use different permission levels for high-level versus detailed reports. Results should also be shown only for groups large enough to preserve anonymity, since combining attributes such as race, gender, disability, and veteran status can identify individuals.
Discussed at 8:47Candidates may skip demographic questions, referrals and internally sourced candidates may not provide the data, and people may distrust how it will be used. The talk recommends explaining the privacy purpose clearly and having recruiting staff consistently ask for the information where appropriate.
Discussed at 12:44Companies are limited in what they can legally ask, and even questions that indirectly reveal prohibited information—such as age—can create problems because intent matters. Required EEOC categories may not match the inclusive categories a company wants to use, so recruiting, legal, and diversity-and-inclusion experts should help design the surveys.
Discussed at 13:19The speaker cites research showing that having only one underrepresented finalist may not improve hiring odds much, while having at least two can make a substantial difference. For a larger applicant pool, she recommends aiming closer to parity with the eligible labor force rather than treating two candidates as a sufficient threshold.
Discussed at 15:05Track counts and pass-through rates by demographic group at every stage, then investigate stages with unusually large gaps. In the example, women dropped off more than men during resume review and onsite interviews, identifying those stages as places to examine and improve.
Discussed at 19:58Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026