Creating an Inclusive Django Community with Kenya Phelps
Published July 15, 2026
This video features JoaquĆn Scocozza at DjangoCon US 2022 in San Diego, California, USA.
Not many people would think that using Django and a PostgreSQL database is a good idea for working with time series data and all its complexity in terms of volume and structure. However, we found out that even the most unusual choices can work, if you have a good reason for doing so. In this talk, we will share our successful experience developing a system with time series data requirements using Django and Timescale, a PostgreSQL based time series database, and the reasons that led us to use this stack. Find out the challenges we faced, pros and cons, and how Django saved the day.
This talk was presented at: https://2022.djangocon.us/talks/working-with-time-series-data-using-and/
LINKS:
Follow JoaquĆn Scocozza š
On Twitter: https://twitter.com/joaquinscocozza
Follow DjangCon US š
https://twitter.com/djangocon
Follow DEFNA š
https://twitter.com/defnado
https://www.defna.org/
JoaquĆn Scocozza explains how his team built a dairy-farm platform that collects sensor telemetry, weather, market, and external data using Django and Timescale. He describes time-series data as timestamped, mostly append-only data arriving in large volumes, and shows how Timescaleās PostgreSQL extension supports it through hypertables, automatic partitioning, hyperfunctions such as time_bucket and gap filling, continuous aggregates, retention policies, and compression. Django models and the ORM handle ordinary application data, while unmanaged models, migrations, and raw SQL are used for Timescale-specific tables and queries; the approach let the team keep both kinds of data together and reach production quickly, although Timescale was not suitable for every situation.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello, everybody. My name is Joaquin. I'm really happy to be here. This is my first time speaking at the Django Con. So we are here today to talk about my experience working with time series data using Django and Timescale. So as I said before, my name is Joaquin. I'm a computer engineer. I worked as a software developer. I talked about Dev, which is a software consulting agency that provides software development for other businesses. So if you want to know more about us, you can check out our website at octobot. io Well I'm from Uruguay, which is a small country in South America
And I'm also a big soccer fan. I really enjoy playing soccer and watching uh different soccer matches Well, let's go with the agenda. So the idea for today is to present the business case context. I will talk about a brief introduction to time series data. I'm going to then talk about the solution we built using Django and Timescale. And finally, some takeaways from our experience. So two years ago, a client came with requirements to build a software product for the daily industry The goal was to create a platform where farmers and managers
could easily access to all the information from the farm. Also, the possibility of monitoring and controlling environmental conditions in the barn with the aim to optimize cows' health and productivity. So uh the platform we built uh gathers telemetry data from different sources. Uh in particular uh information from barns and its devices, weather information , markets, and also external integrations that provide even more data from the farms and its barns. So how does this data look like?
One of the first things we tried. uh was uh experimenting with the the data. So here we have an example of how uh sensors telemetry data look like So basically we have a value which represents the reading of that sensor and also a timestamp which is the time associated to that reading. Mostly all the data we gather looks like this, it's basically a value with a timestamp So after we looked at this, some questions came up. In particular, we were wondering if, for example, we could use the traditional SQL. use
two or do we need non-relational databases or can we use our preferred stack which is React and Django Also some concepts came up like real-time systems, time series, and also the necessity from the clients to go fast to market. So the next thing we are going to talk about is time series data. So what is time series data? And here we have a concept that I really like which is uh basically it's data that collectively represents how a system, process or behavior uh changes over time. So it's basically that
Some characteristics of this type of data is, for example, it's time-centric, which means that all records have a timestamp. which also means that time is a primary component that helps uh that helps us to analyze uh our data and divide meaningful insights It's usually append -only , which means that we usually want to insert new data because it usually doesn't change its uh It's also recent, we could say it's recent because new data is usually about recent time intervals As I said before, we really we really want to make updates or gap filling
from old data. Here we have an example, this is a screenshot from our solution. a chart which represents uh different uh which represents the temperature uh in the last week uh every one hour We have the time represented in the x-axis and the temperature on the y-axis. For time series data, we usually we always have a time as the x-axis. It's the primary access for this type of data. And there are lots of examples. For example, we could also mention stock prices, health rate monitoring. uh
the monthly subscribers for a blog, let's say, and many more. If you think about it, time series data is almost everywhere. So one of the first questions that came up is, well, can I use a traditional database for working with this type of data? The answer is yes, but that's not something I would recommend because time series data is different in several ways For example, time series data is usually ingested in massive volumes, which means you will receive lots of data when you are ingesting this type of data. Also the life cycle of this data is important
once you start storing data You need to care about what to do with old data if you don't want to run into storage issues issues, for example. And also because uh you usually want to run um large scans of data uh when performing queries, for example So because of this, there are some special databases called time series databases, which as the names Sorry, uh that are optimized for time series data and they have key architectural design properties that make them different from other databases
and very useful for working with this type of data. Some examples are Influx DB, Apache Druid, and Timescale, which is the one we use, and many more. So the solution we built is built on top of Django and Timescale So the first question is what is timescale db? It's a relational database for working with time series data It's also an extension built on top of Postgres SQL, which is actually pretty good. So why did we choose uh timescale?
Um well first of all, and probably the most important is because it's time series oriented, which means that has architectural uh key properties that make it special for working with this type of data and also because it has specific time series functions that help us work with this type of data easily It's so because it's open source um because uh as it is built on top of Postgres SQL we can leverage the full SQL language. Also because it has a cloud version, which we actually use, and because it has many resources like documentation, Slack community, etc.
So, how do we get started with Django and Timescape? The third thing you would like to do is to configure the settings for working with this database. And as it is built on top of PostgreSQL, we we can actually use the same configuration we would use for a PostgreSQL database. So here in this example we are actually using the same driver as we would use for PostgreSQL and then we would just need to define the credentials we need for our database So how is time series data stored in timescale? Timescale defines what are called hypertables, which behind the scenes are a group of Postgres
SQL tables called chunks. So the idea is that each chunk in the hypertable stores data for a particular time period. So in this example, the hypertable on the right stores information for a particular day, so each chunk uh represents uh a day in the data So what timescale also does is automatically partition data for us so we don't have to care about partitioning data when inserting data and this architecture provides uh improved performance and user experience when working with time
series data Timescale also creates indexes for us in each of the chunks. So that also helps to improve the performance. So the next question is uh how do we work with hypertables in Django? So to explain to explain this, um I think it's better to do it with a with an example So let's say we want to store temperature readings from a sensor or from different sensors into a hypertable. So the first thing we would need is to create a tango model to store the different sensors. And then we would need a hyper table with for example these columns which are value, time, device or sensor ID.
and a primary key to ensure a primary key with the time and device ID to ensure uniqueness of time and device ID So is it plug and play? Is it that easy? Not so fast because uh one of the things that Django currently does not support is uh composite primary keys that the like the one we would like to use here. So there are different approaches. Uh we saw on the internet for example, One would be to install a third-party library that easily handles the uh Django and timescale settings and integration.
By the time we did this solution, there were not or we didn't see any third-party libraries mature enough that we could use. Also, we saw some dark solutions like editing data migrations for the primary key, for example. So finally what we did is we f we used the unmanaged models from Django. This is how the data migration would look like for creating our hyper table. We basically have to create the table and then define the primary key we want and the last step is to use a create hyper table function from timescale. Here you say okay this is my table and this is the time column for my hypertable
Easy as that. And then we will have an unmanaged model for these temperature readings. And notice that you have to set manage folds and specify the database table. So what happens with the other tables? The truth is that not everything is a hyper table in timescale. So we use traditional Django models for working with non-time series data, which is actually great because uh with this uh With this approach, we were able to have time series data and also application data in the same place So how do we query time
series data with Django? The truth is that we mainly use RAW SQL queries using time scale hyperfunctions for querying time series data Hyperfunctions are a special specialized set of functions that allows you to analyze time series data. One quick comment here is that when using RAW SQL you have to be careful with security uh stuff. For example, you need to you would need to do input validation to prevent SQL injections and many other things. And then we use the Django, the power of Django's ORM for non-complex queries and inserts.
Okay, so let's say we want a query Barnometrics. In this screenshot we have a customized dashboard that users from the app will customize So um a very simplistic query of this type of data would look like this in this case We are trying to get a five-minute average of temperature data. So here we are using one of the timescale hyperfunctions called time bucket, which we use a lot in our queries. In this case, we are using TimeBucket to aggregate data every five minutes. TimeBucket works pretty similar to the date
trunk function from Postgres SQL. With the difference that it's more flexible in some ways, we can use arbitrarily size time periods Like in this case. So we basically include time bucket, we say okay, we want the average for the temperature column, and we group that uh by uh Our bucket. So another example of this type of query would be using the time bucket gap field, which is actually pretty similar to the time bucket hyper function In this case, we are trying to get a one-day average of humidity data
for the last week. So time bucket, gap fill. Will automatically fill those gaps in which we don't have data. So let's say that from the last week there are two days in which we don't have data So this hyper function will automatically fill those gaps for us. In this case, since we are not specifying any value for those gaps, the the query will return return none for these values. This is actually really useful for uh working with the graphs or when creating charts. So we want uh every day to have a
An associated value. Another example, and this is also another screenshot from our application, is the Windows charts For those who don't know or are not familiar with the Windrow Charts, it basically tells us the wind distribution for a period of time So here is a query we used and here the main thing to notice is the Hisogram hyper function. which really fits in this context because it provides us with how data is distributed In this case for a period of time, you can specify the minimum and max uh
ranges for our data and it will provide us with um the distribution of wind speed Then another interesting concept is the continuous aggregate. It's designed to make queries on ver on very large datasets run faster. So it works pretty similar to the Postgres SQL materialized views. But with the difference that it can be continuously and incrementally refreshed. So we can define um Refresh policies to automatically update our materialized views in timescale
And also timescale handles the concept of real-time aggregation. So usually our materialized views don't don't include all recent data uh and let's say we want to query data from uh today Timescale we use what's uh what we have in the materialized view plus uh the recent data So if we want to do this in Django, we could use a data migration to create the materialized view. Let's say we want the daily average of temperature We would create the materialized view and then we would have a continuous
aggregate policy to refresh this materialized view every one day. Another concept that is related to the aggregation, to the continuous aggregate, is the one of downsampling So when data gets solder, you usually don't really care about individual records of raw data. You usually care about how aggregated data looks like in time. So one thing that you could do is define retention policies to start dropping old raw data. and saving storage
space. So here we have an example of how our attention policy would look like in Django It's basically you say, okay, I want this hyper table to have this retention policy and in this case uh it's defined uh with an interval of ninety days, which means that after ninety days our data will be dropped. And then finally but not less important , actually it's one of the most important concepts in timescale, is the concept of compression. Which is related to the data lifecycle management. Timescale offers the possibility of compressing our data
So let's say we have this hyper table with these columns. When compression is applied to this table Timescale will compress all the rows into one , keeping the values into arrays for each column This reduces the amount of storage required in our database. It also increases the speed of some queries, for example, those that are particular for that row and one uh one thing to bear in mind is that compressed chunks cannot be modified so usually you want you want to compress all the data
Because recent data is usually being modified. If you want to modify a compressed chunk, you would need to manually decompress that chunk And here is how we would define a compression policy with Django. We would need a data migration. You say, okay, I want this hypertable to have compression enabled. I want compression to happen after three months and that's mainly it. Then timescale will run jobs in background that will start compressing our data based on the compression policies we have. So takeaways. From our perspective,
Django and Timescale was a very good solution for our case In particular for working with time series data. It allowed us to have data in one place. As I said before, we had time series and application data in one place We were able to leverage all the knowledge in Django and SQL that we already had, allowing us to go fast to market. And also we were able to use the cloud version of timescale, which was really useful because it provides us with metrics And dashboards to easily manage our data. And another conclusion we had is that timescale is not magical.
The truth is that We had uh in some cases we were not able to use timescale. In some situations we would have liked to use it, but uh In the end it was pretty helpful. So here we have some stats about our solution. Currently we have more than 80 farms in production, more than 300 users. more than 40 hyperchanges and more than 300 gigabytes of data and more than two years of development and counting. So I hope you have liked this presentation and thanks everyone for joining.
Time series data represents how a system, process, or behavior changes over time. Its records have timestamps, are usually append-only, and are commonly based on recent observations.
Discussed at 3:27Yes, but the speaker does not generally recommend it because time series workloads involve high ingestion volumes, lifecycle management for old data, and large scans during queries. Specialized time-series databases are optimized for these patterns.
Discussed at 5:44TimescaleDB is a PostgreSQL extension and relational database designed for time series data. The team chose it for its time-series architecture and functions, PostgreSQL and SQL compatibility, open-source ecosystem, and cloud offering.
Discussed at 7:17Because TimescaleDB is built on PostgreSQL, Django uses the normal PostgreSQL database configuration and driver. You provide the usual database credentials and connection settings.
Discussed at 8:49TimescaleDB stores time series data in hypertables, which are managed groups of PostgreSQL tables called chunks. Each chunk covers a time period, while TimescaleDB automatically partitions the data and creates indexes to improve performance.
Discussed at 9:36Django did not support the desired composite primary key directly, and the available third-party options were not mature enough for the project. The team used unmanaged Django models, created the hypertable and composite key in a data migration, and then registered the table with TimescaleDB.
Discussed at 11:56The team primarily uses raw SQL with TimescaleDB hyperfunctions for time-series queries, while using Djangoās ORM for simpler queries and inserts. Raw SQL requires normal precautions such as input validation to prevent SQL injection.
Discussed at 13:31time_bucket groups measurements into configurable time intervals, such as five-minute averages. time_bucket_gapfill also supplies missing intervals, which is useful when charts need one value for every time period.
Discussed at 14:18Continuous aggregates are similar to PostgreSQL materialized views, but they can be refreshed continuously and incrementally. Refresh policies update them automatically, and TimescaleDB can combine the aggregate with recent data for real-time results.
Discussed at 17:27You can downsample older data into aggregates, apply retention policies to drop raw records after a defined period, and compress older chunks to reduce storage. Compressed chunks cannot normally be modified, so recent data should remain uncompressed until it is no longer changing.
Discussed at 19:03For this project, it was a good fit: it kept time series and application data together, reused the teamās Django and SQL knowledge, and helped them reach production quickly. The speaker also notes that TimescaleDB is not a universal solution and was not suitable in every situation.
Discussed at 22:07Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026