Creating an Inclusive Django Community with Kenya Phelps
Published July 15, 2026
This video features Christopher Clarke at DjangoCon US 2014 in Portland, Oregon, USA.
By, Christopher Clarke
Developing applications for civil service organizations can be uniquely challenging. The presentation discusses our experience with MASS a market surveillance and monitoring application developed for the TTSEC. MASS is based on Django and Pandas. We highlight not only the techinal aspects of the solution but also address the HR and organizational factors that impacted on the project
Help us caption & translate this video!
Christopher Clarke explains how his team built a web-based stock-market regulation and surveillance system for Trinidad and Tobago after a new trading platform failed to provide the regulatory data and monitoring tools they needed. The system uses Django with PostgreSQL, Nginx, Gunicorn, lxml, web scrapers, and Pandas to ingest finalized trading data, reconstruct links between orders and trades, calculate statistics, and provide dashboards, filters, pivot tables, and exports. He describes the practical constraints of government software projects, including slow procurement, changing staff, poorly understood fixed requirements, legacy hardware, and outdated operating-system packages. Pandas is used for data cleaning, time-series and statistical analysis, dynamic pivot tables, and rendering query results as JSON, CSV, HTML, or Excel. An IPython Notebook integration eventually let analysts perform ad hoc analysis themselves, while the Django Pandas package moved data-access and business logic into reusable managers, making the views much smaller and easier to maintain.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
So I'm also Real Chris Chris Dev on uh Twitter. And I work for a small development company based in Trinidad called Chris Dev. My talk today describes the challenges that I faced in building a stock market regulation, surveillance and monitoring application, and how we overcame those challenges. Using a combination of Django and Pandas. I don't work for the Securities and Exchange Commission of Trinidad and Tobago, so my views are my own. Right. So our talk is divided it uh up into four sections.
So the first thing we do is uh Talk about some background to the project, then we'll introduce our solution architecture and then we'll look at our Python or Django reusable app called Django Pandas And then we take your questions, of course. Okay, so in 2011, um the Trinidad Island Tobago stock market was forced to migrate to a new trading platform. This was because that the company that made the trading platform had gone bankrupt some years ago and could no longer um maintain the platform or could no longer were given support on on the platform.
So The new platform was great for traders, but didn't uh provide a lot of the regulatory hooks. That regulators needed for monitoring and surveilling the stock market. I had done some previous work with the Securities and Exchange Commission, so I was asked to present a proposal for a system that could bridge the gap. The requirements weren't really well structured, but the two things that they wanted was that the system should be web-based. Because they wanted to share whatever system was developed with other regulators, such as the
central bank, Ministry of Finance, and so on. And the data we must or the system would would use would be complete and finalized data. I never actually found out what that meant. But hey. Um and they said if our proposal was accepted, we would then work with key users to develop a set of more detailed requirements for the project. Well of course that is a uh uh could only lead to one thing The first thing we encountered was a very slow procurement process. I mean, we expected that as a government.
kind of bureaucratic organization. But six months between tendering the proposal and discussing and accepting and sign off to the project that's uh little long but the more uh actually more surprising was the fact that um There's a lot of mobility. One might think well being a people working bureaucracies, they're there for life But that's not really true. There's a lot of interdepartment mobility, people m leave and so on. So that um these so-called key users um to to tell you truth Um, there's only one left who's uh who by the time we finish the project and we deliver the system
to the key users. Who framed the requirements in the first place. They've all gone on, they've been promoted, they've left the organization and for so forth. And what we found is that a lot of the people replacing them You don't have the kind of statistical training or technical background for this kind of system. And you in any case you have to try to um uh uh explain, promote, sell the system to people like a new set of people every three or four months, which is leads to a number of problems, you know, you have long lead times of decision making and the and then if the requirements that we had
um They were not well understood by the the final users because the people who specify in the requirements, they're no longer there And the people who so the people are using so why do we have feature X? Because that's a requirement. Why is that a requirement? So you have to go explain the rational, sell the project once again So um and of course people will say well oh we should you should use agile methods you know um to deal with such problems but In government contracting, um you find that there's an implicit waterfall model, you know, so that
And once the requirements are put in place, they they need to stay there for the until the end of the project. Even though they They're no longer relevant or nobody understands why they're there. So hi. Um so uh despite these challenges we were able to deliver system to the end users. Um we started ruling it out. in the third quarter of uh 2013. It's called Mass Market Analysis Analytics and Surveillance System. And it's a pretty, it's a standard Django application with a couple of exceptions. Um, so this is this is the basic architecture components here.
The things that The front is obviously we why why are you using pandas and then you see that Red Hat Enterprise and X5. 2 and Postgres 8. 2 We'll well in in subsequent slides we'll explain why. Here's the system diagrammatically. So the setup the the our uh and uh maybe I've went through the slides too quickly, but though we have some really ancient Dell machines. uh as silvers but but that's a uh a legacy of uh silver consolidation exercise that was done in 2010
and again in government style organization. There's a five-year period. So whatever is done, that has to stay. And we can't upgrade till 2016 How obviously to build a modern Django application because Red Hat Enterprise Linux comes with Python 2. 5. We had to upgrade um Python and and we are in the proc process of migrating the system to Postgres 9. 3. It we we if you know in the diagram we we had uh we have two servers. Well we are uh we originally had three, but one was taken away from us because
an exchange server failed so it took our server away and we had to do with these two silvers. So um we have uh two Goni corn um instances and we use uh Nginx um um wrong robin load balancing um The PG pool is because we were using um um because we are using Postgres 8. 3 or 8. 1 when we started is required for replication. So we replicate the database between the two servers so we get some kind of redundancy or
And so when we when we finally migrate to Postgres 9. 3, we might get rid of the PG pool and replace it with something like PG Bongsa. Um the system itself is very read centric um because um So we only have one place that we really upload data. It's not a real-time system because again The the users or the client um wants to only look at final data. So we in terms of data, we we retrieve data from
the trade and systems um servers log. We actually have a request. FTP or SFTP sister job that goes up and retrieves this log file which is our XML monstrosity and we parse it using uh and and and load the data um thanks to lxml because you know it's it's really a a horribly formatted Windows file. And so aside from the trading data, we also have a bunch of scrapy spiders that we used to get
um data from various stock markets in the region because they have to use it for comparisons and that sort of thing. And okay, so what is pandas used for in the system? So we we obviously we do a lot of statistical and time series calculation, so there's a four tier of pandas And then what what one of the one of the things that we found that we needed to do because it wasn't implicit in the requirements is that we needed to link trades and orders and uh be able to um rebuild the history of uh trade from when it the order was originally placed
So when it when it was changed and and and so forth. So we we are put we use um Panda 's um approach of Split apart apply combine to reconstruct trade order books and calculate duration and duration statistics. An interesting use of pandas in the system is the ability to do pivot table analysis on the fly. So we actually generate pivot tables on the fly. Pandas is where has a lot of facility for data cleaning and we actually have use pandas in our um Django views
because it could be easily rendered so a pandas data frame could easily be rendered as JSON or CSV or HTML or Excel. And we have a project called Django Pandas, which provides a manager and some other functionality so that we could easily render pandas data frames from query sets. So let's talk a little bit about denormalization and and cassion. So with a system that is heavily statistical, with a lot of statistical calculations, obviously, we don't want to do that on the fly. So we have uh daily summary data.
Um uh so tables or models for storing daily summary data and these are populated when we actually load the the raw data. And we're also gonna extend the strategy to also cash at the database level, if you want, the more expensive statistical calculations. Um strangely enough we're not really making that extensive use of caching uh um except at the level of the dashboard um uh however we we we intend to use um Django cache machine to Um
do some query set caching. Um and uh The w the the the the thing that is holding back both um data denormalization and caching is the fact that the user requirement is that they must have Total flexibility over date ranges and security. So there's no standard um queries. So you can't, you know, so you 're kind of limited at as to what you know what what you could cache because they want to put in any data range to so that they can get the statistics calculated. We're working with them to change their minds about that that requirement. So this is uh this is these
these are some screenshots for the mass web app. So this is the dashboard showing some of the statistics calculated This is our activity filter which allows the users to drill down from a trade and they could click on it, any trade, and see all the order history. behind the trade. And uh now you'd expect this this should be easy because clearly in any real system if you have a trade they would have a key linking trades to orders But for some strange reason in this trading system, there was no link between trades and orders. So you actually have to use Panda 's capability of getting a bunch of candidate groups.
where that where this this this uh this trade could be linked to and then find the actual trade so that uh I don't know if you could see uh there's a green thing there so that's the one that is the the trade uh up above so I don't think that could have been done easily without pandas and and it's only about four or five lines in in pandas. or using panel. So this is a pivot table. So in this case, um we take in a trader and We looking at as trades over a particular period and we pivoting um or we breaking down by security and client and we
looking at the number of trades and you can drill down from this into the um to to find out more about the or the trades and orders that made up this this any cell in the table. Uh so that this was we use Foundation Five and um the It it's um it works best in Chrome and the users love that because they don't they hate IE for some reason. Um right not only developers hate IE but users do even corporate environments so um so high charts is useful chart and high charts and high stock
um We use jQuery data tables for the interactive grid. The only criticism about that, about jQuery data, is not responsive. But it it's it's it's improved a lot. Okay, so that that's the traditional client if you if you want to call it that. We also use the iPython notebook with this project The idea came from one of our directors who wanted to to give the ability of the analyst to to perform ad hoc statistical calculations using the data that we had in the system.
So uh we we tried it, we did a number of of experiments and then we came up with the best way to do it was to basically install Anaconda Python on the users Windows workstations Put the the actual project source code and we have a minimal settings module. And where the I the installed app only points to the to the to the to the models that you want to expose at the iPython notebook level And we create a profile for each application that we want to use.
And I f iPython has this um startup directory. So you put a script there that gives you that sets the part to include your application and sets the Django settings module so then you could import the Python object in the context of a iPython notebook. Or sorry, the Django objects in the context of an iPython notebook. So then We can just start up a notebook using the command and then to make it even easier we create a batch file and stick it on the desktop so they could So they could start and they believe it's kinda easy. Right?
So um what was the in user 's initial Reaction, oh, you want to turn me into a programmer? No way, right? Okay, so yes, yes. But a few weeks later, um The regulator, the central bank of um wanted historical data from mutual funds And they wanted it in a certain way. They wanted for each issue, they want all the funds for that issue. They want one spreadsheet, one workbook and a spread uh a sheet for each. is here and that's like hundreds of it's about 40
something unit trust um um issuers and hundreds of different funds and they they were going to do it manually. So one of the analysts says and they wanted it tomorrow yeah So one of the analysts said, I I I I heard you blabbing on about something that this thing could produce uh Excel spreadsheets. Can you do something about that? So I sat down with the analyst And we came up with this, which basically did what they wanted. Right? So when it when they saw that, they became enthusiastic. Right? And I I put together some wrapper functions and um I added some documentation to the notebook.
So this is a section of this user's iPython notebook that they use all the time and they've been able to modify it. I just added the the wrapper functions and I I basically take some of the more complex stuff and I put it in the startup file So, right. And now I uh I'll talk a little bit about our Django Pandas module and it It is very it's is in it's it's been useful in our in this project because we don't want to use Django in the traditional way, but we try to Use Django to the iPython notebook. So it has two basic modules, an I.
O. module and a data frame mod uh mod. So this is a typical Django model Um this is uh how you use the read data frame method, which is currently the only method in the I. O. module. So it could uh you could create a data frame with all fields specific fields, you can give you can it create an index and you could use filters and excludes The data frame manager is based on the Past True Manager by PM McClannon and it's in Django Model Utils, which is a great package by Carl G.
So what you do is you overwrite your default manager with the Django, the data frame manager, and that this gives you access to three methods. Two data frame, two time series, and two two pivot table. The two data frame meta does exactly what the read IO. The two tra time series method supports two kinds of storage models for time series long and wide. So the wide is basically a bunch each column in the off represents a A column in the time series. A time series is a data frame indexed by a daytime object.
In the long time series model, each um The the the the so if in this structure there's three different series and it's stored long. And when you use the two time series methods, it pulls them out and structures them as a data frame with a time series index. The two pip we also um uh as I said before support but the real the real thing about it is um Yes, you could use the but the real uh is the is to um subclass the data frame
manager and Build your own custom manager with the methods and business logic that you need. So when I started , When I started my my views used to look like this. I had a a mix-in with all the business logic to create a trade data frame. So my view classes were huge. And then I realized I could just subclass my data frame manager and then I could put all the logic there so then I just have uh a single line of clue that that does you know what what I had before.
Right.
Pandas is used for statistical and time-series calculations, rebuilding trade and order histories, generating pivot tables, cleaning data, and rendering query results as JSON, CSV, HTML, or Excel.
Discussed at 10:35The application precomputes daily summary tables when raw data is loaded and plans to cache expensive calculations at the database level. Flexible user-selected date ranges make caching harder because there are few standard queries.
Discussed at 12:54The system uses Pandas to identify candidate groups of orders for each trade and determine which group actually corresponds to it, allowing users to inspect the order history behind a trade.
Discussed at 15:14The project exposes selected Django models through IPython Notebook, using Anaconda, a minimal settings module, and startup scripts that configure the application path and Django settings. Desktop batch files make launching the notebooks easier for users.
Discussed at 17:33Django Pandas provides I/O and data-frame functionality for creating Pandas data frames from Django querysets, including filtering and indexing. Its data-frame manager also supports conversion to data frames, time series, and pivot tables, and can be subclassed to hold application-specific business logic.
Discussed at 21:26Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026