Python for Planet Earth: Climate Modeling and Sustainability in Action with Drishti Jain
Published October 23, 2025
This video features Drishti Jain at DjangoCon US 2021 in Online.
Have you ever thought of using data visualization to represent data; but feel that it is a cumbersome process? Worry not β Orange is here to the rescue!
Come, dive into the world of this magical open-source data mining tool that can also be used as a Python library.
Beginner-friendly!
This talk was presented at: https://2021.djangocon.us/talks/illuminate-data-with-visualization/
LINKS:
Follow Drishti Jain π
On Twitter: https://twitter.com/drishtijjain
Website: https://linkedin.com/in/jaindrishti
Follow DjangCon US π
https://twitter.com/djangocon
Follow DEFNA π
https://twitter.com/defnado
https://www.defna.org/
Video production by the speaker and DjangoCon US 2021 Volunteers.
Drishti Jain explains how data mining turns raw data into useful knowledge through retrieval, preprocessing, analysis, and interpretation. She presents data visualization as a way to reveal patterns, relationships, trends, and anomalies more quickly than tables, while emphasizing that good visualizations need relevant, clean data, a clear goal, appropriate visual forms, and domain knowledge. She distinguishes descriptive, diagnostic, predictive, and prescriptive analytics; outlines the volume, velocity, and variety challenges of big data; and demonstrates Orange, an open-source, visual, component-based data-mining tool that supports interactive workflows, machine-learning methods, and intelligent visualization such as selecting the best number of k-means clusters. She argues that combining computational tools with human judgment improves decision-making, storytelling, data literacy, and awareness of bias.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hi all and welcome to today's presentation. So there's a ton of data that we see around us and have you ever thought how do you make the most sense out of the data that is there around you How do you eliminate your data using visualization so that you are able to see things which are beyond just the obvious? Hi everyone, I'm Drishti Jan and today I will be presenting this talk on how do you illuminate your data with visualization. How do you make the most of the data that you have and use data visualization as a key tool in order to identify patterns, relationships, information from the data that you have gathered To tell you a little about myself, I am a computer engineer.
I develop software on a day-to-day basis. I'm a tech geek. I'm a social entrepreneur and I have a non-profit organization that works in 11 cities across India in the sectors of environment, education and healthcare. Wherein I help the underprivileged children use technology, learn about technology and one day be able to achieve their dreams so that they have the necessary resources which they might otherwise not have. I'm also an international tech speaker. I love to give back to the community, speak. I often travel across the globe speaking at various technical conferences. Presenting a across a variety of subjects, talk about various open source projects that are out there, and I'm also a mentor to a number of uh students
who are pursuing their engineering or their bachelor's degree in computer science related fees so that they're able to um research in the particular area that they are that they aim to research in or get the job that they want to get it get the scholarship the mentorship and everything So when we talk about data, there's a lot of things that come into our mind. Data can be very sparse. Data is there. I know what is the obvious interpretation of it, but I am not able to Actually understand and identify new things which might not be very obvious. So that is where determining comes into place Data mining is an autonomous process of discovering previously unknown patterns. We are trying to identify new patterns which are useful, novel, and um
Understandable from large data sets. Now, this is very important because uh getting knowledge about things which we did not know earlier is very very crucial to understanding the various trends and patterns in the data that we have. Data mining. uh has four uh particular steps that take place in it the very first one being data retrieval so Data retrieval consists of gathering all the data. This can be through uh this can be data which is present in various forms, structured, unstructured, quasi-structured. Retrieving data from various sources, accumulating all of them. The next step that comes is data pre-processing. So now that we have all of the data from a previous step But we cannot just start finding patterns on the data that we have collected.
We want to pre-process the data so that we make more sense out of it. so that we are able to process data uh uh handle any missing values handle sparse data uh fill in those values with the correct logic of particular values for example it could be mean median mode based on the technique you are using And create a good data set because having a good clean data set is crucial to doing anything further in the process The third one is data analysis. So data analysis is when we interpret and try and make patterns, relationship, identify trends in the data that is pre-processed. So this can be if
uh the domain that we're working in is where the prediction so this could be wherein we are able to identify trends of how global warming s global warming is affecting the way seasons and the duration is reducing or increasing, when is it likely to be uh when is it that we can experience heavy rains? When is it that uh there could be a situation of a drought in uh uh in various places across the globe. All of this comes under data analysis. The last but not the least is interpretation. So now that we know that this particular trend is what we are observing So how do we use it to our advantage? In case of uh weather prediction, how do we use this information? So for example if we have identified a flood could come in a particular area.
So we can evacuate people who are present In that particular area during that time of the year, or issue warnings so that we know when the severity gets high. Before the severity gets high, we know when to evacuate people from the particular area so that There's no loss to life or property. All of this comes under the process of data mining. Now the third point that I was discussing previously was about data analytics and that is very crucial to identifying trends and making predictions, getting to know how do we optimally use data. So there are four types of data analytics. The very first one being descriptive. So what is exactly happening in my business?
This is wherein you make interpretation based on the data that you have You are trying to identify information from the data that you have accumulated. How do you make more sense out of it? That is what comes under descriptive type of data. The diagnostic type of data is why is it happening? You want to know a particular reason. Taking an example of a financial institution and you're trying to identify customer spending habits. So you identify that a few months of the year have more uh customer transactions, more spending by customers So why is it that it is happening? Is it some festival? Is it the holiday season? What exactly is influencing it? So knowing that, identifying that, diagnosing it
comes under diagnostic data analytics. The third one is predictive data analytics. So what is likely to happen For example, if you pick up a particular month, say May of a particular year, so what is the likely trend or pattern of customer spending habit that can happen at that time So, what is likely to happen based on the previous results that you have been working on? That comes under predictive kind of data analytics. This can also be useful for Predicting things way in advance based on the analysis done on the previous data. The fourth kind of data analytics is the prescriptive type of data analytics. Now that you have the information, you know what is about to happen.
So what can you do? If you know that a particular month is going to be um a time wherein people are ordering more things At that time you as a company can be prepared with more inventory so that you do not have delayed shipping times. So that comes under prescriptive. You have interpreted uh Based on the data that you have, you have the information, you have the knowledge. So, how do you make use of it? You even if you have it and you do not make changes or do not make use of it, then it is of no use. So prescriptive type of data analytics works in that domain. Now many a times uh many people get confused between what exactly is business intelligence, what exactly is data science, when I'm working on
particular thing, what exactly which category does it fall under? So business intelligence is whenever you're trying to understand what has happened What is the way that the company was performing in the last quarter? How many units did a particular company sell? All of these questions get answered in business intelligence. You are trying to create reports, analyze the way the uh company has been performing all of that comes under business intelligence. Data science is more about how do you use this data to predict things Even if it is not the perfect prediction, but if it is an optimum prediction, it is very useful. So what if This particular thing happened? How will it impact the business? And what is the optimal approach towards taking our business forward
All of these questions get answered in data science. It is more about the analytical approach, it's more exploratory. You're trying to fit in different features, different situations and see what the outcome would be. Now that we have been talking about data, uh there are a number of challenges that come with data, and the reason is that The data that we deal with nowadays is all big data. Big data is nothing but numerous amount of data from a number of sources generated at a very high speed very frequently and we need to handle that data such that we are able to make the most out of that data. The three key challenges that come when working with big data are volume, velocity and variety.
The very first one being volume. So the scale at which data is generated. Consider a case. Consider your daily life from the time you wake up to the time you sleep And actually even when you're sleeping, if you're using like a fitness tracker, you're tracking your sleep and uh was it a sound sleep or not? All of that is tracking data. It is creating data. So the volume at which data is created during your day, you will be using your laptop, your mobile phone, your iPad, your smartwatch, your Um smart refrigerator, smart lights, any and everything is generating data. So there's a lot of data that is generated on an everyday basis by each and every person. The second is velocity.
The frequency, the speed at which data is generated is very very huge. Because of the number of devices and the way we multitask many times it leads to data being generated at a very high speed. So handling that data uh storing the particular data that is incoming, making space for the new data that is coming, handling that high velocity of data is very critical when working with big data And the next one is variety. So there are different forms of data. It could be structured, unstructured, quasi-structured data. Consider a case wherein you are on any of your social media applications. At that time you're just scrolling through your device and you are actually generating a clickstream kind of a data.
You're clicking on one thing, you're going to a particular page, you're liking somebody's photo, you're commenting on a post uh in all of these things you're creating a lot of data also uh while scrolling itself you are generating data so the post wherein you stop uh what is the screen time of a particular post All of that is also data which is captured by social media uh applications in order to tailor advertisements, marketing towards you as a customer. So that these are the critical challenges which have to be handled when dealing with big data. Now talking about all of this, uh data visualization is a key tool which comes in very very handy and very useful whenever we are working with large amount of data.
Data visualization is nothing but a pictorical representation of the data that we have The reason that data visualization as a tool is very effective is because anything that is visual tends to give us a better understanding of what the data is trying to tell us. Taking an example of a case wherein we are trying to map which places are colder and which places are hotter, what is the average rise in temperatures across the globe, across continents. Having it visually represented through colours is much more efficient as compared to just a table with the countries and the corresponding rise in temperatures. That will make it difficult for you to interpret that data.
Whereas once it is visually displayed to you, it is very easy, very quick to interpret data Data visualization is also considered a modern equivalent of visual communication because it is trying to communicate to you what is the information the data is trying to share with you. Also, data visualization helps us merge data from different sources. So, having one data visualization and imposing one more data visualization on top of it will help you understand things much much better as compared to not having. having uh data visualized. One of the key things when doing data visualizations is the use of colors
So having various things represented in the form of colors is much more effective as compared to just using grayscale or black and white. Using more colours will help you discover trends easily. When you're using data visualizations, you're plotting things. You are uh giving a visual view a pictorical view to the data that is there that will help you see trends and that will help you discover trends in anything and everything that you're doing You will be able to comprehend information quickly. For example, if you want to know if there has been a rise in customer sales or not. Pro plotting it as in a line graph format will give you the answer to this question within a matter of seconds
Whereas if you were to manually go and see the sales in each and every month of the year, that will take some time for you to understand. Plotting it, having a data visualization of the data will help you comprehend the information very very quickly. And an interesting thing, it will also help you identify the relationship and patterns between these things That will actually help you in interpreting the data that you have. What exactly can you make out of it? Or how has the data been performing? Is it the way that the company and the business model has been expecting it to do or is it something else? So all of that um Is achieved through data visualization. And it is all about finding the right balance.
Just using a lot of complex mathematical models on data will not be useful if we are not giving it the human touch. If the understanding of the domain that we have, if we do not have that, in that case it will not make any sense. For example, if we know that the season is winters and at that time summerwear clothes will not be purchased a lot. So having this domain knowledge will be useful in understanding and applying the right mathematical model for the data so that the correct information is uh being used to find Future predictions using the data that we have. So let me just show you how
impactful data visualization is. Let's consider the cybersecurity breaches across the US. So all of this information is here in a fabulous format. It has thousands of rows, and this is just a snapshot of a part of the table. Take a second and see. Is it difficult to interpret the way cybersecurity reaches across say the states of the US have been? It is right even in thousands of entries whereas in case of big data we deal with millions and billions of data points So was this information difficult to read, interpret? And did you find the information being conveyed through the data useful?
It is questionable because you were not able to quickly understand what the data was trying to tell you. Now looking at a visual plot of it So here is a mosaic display of the number of individuals affected by the cybersecurity breaches across the US as per their state. And it is also color-coded. So this is so much more useful. You are able to identify what exactly is the data trying to tell you. You are able to identify a relationship. Which state had more individuals affected? and in each of the ranges of individual affected, which state had a major role in it. So you can use this information for something useful. If you're a um
Antivirus company or if you specialize in stopping any cyber attacks, you know which state to target with because that would have more sales of your particular product So you see how data visualization can be very very impactful in all parts of life And there are four key things that makes a visualization to be a good visualization. The first one is information. Having very less data will not give you the desired results. So having the right amount of data and having it uh to be noise-free and uh to have good values, relevant values is very important. Having the story or concept means Is it of the particular domain and does it make sense?
Just having random points will not uh be very useful. Having it relevant to the use case you are trying to solve is important. You should know the goal that you have. What is the ultimate aim? Are you trying to boost your sales? Are you trying to reduce your costs? What is the ultimate goal? Based on that, you'll be working on the data and using data visualization. And using the right visual form, using the right uh data visualization technique. Do you have to use uh pie charts for earnings or do you want to use histograms for earnings? Is it about sales It depends on what you're trying to solve. So having the correct form of data visualization for the correct situation is very crucial. We've been talking about the various advantages of data visualization.
A key major advantage of data visualization is you can illustrate, highlight, or hide Data points which are not relevant to your case. Consider in this case of highlighting particular data points. So some data points will be influencing the whole data. So you can highlight those and concentrate on those. That can be easily achieved through data. Visualization and if there are data points which are irrelevant to what you are doing, they're just anomalies which happen once in a while. In that case, they can be suppressed, they can be hidden So data visualization can come in handy in a number of ways. And when you're highlighting data, you can see data and context as well. In what two particular features are you is that more relevant?
You can have a 3D mapping of data and understand what affects more. How do you discover trends? Having it in a having data being represented in a visual format will increase your chances of exactly understanding the hidden information through data by identifying unique uh previously unknown patterns. The big data ecosystem is pretty big and uh Pretty complex, but dividing it into four key categories. The big data ecosystem has the data devices. This could be your video game controls, your ATM cards, card readers, RFIDs, computers, cell phones. All of these are data devices, these are data generating devices.
The second one are the data collectors. So when you are shopping online , when You're calling people, uh, you are subscribing to a particular service on your mobile phone with your carrier. All of these people come under data collectors. They're collecting information about you, about the way you interact with things The third are the data aggregators. So these are websites, information brokers, advertising analytical services who are trying to aggregate the data which you are generating through various sources. There's a particular data that you are generating when you're shopping online and a different data that you are generating when you are using some add-on service on your mobile phone. So aggregating all of this is the work of data aggregators. The fourth one are the data users and buyers
So these are the institutions and companies and individuals who are using the data that has been aggregated by the data aggregators to make sense and use it in their business. So these could be media firms, banks, if they in case of banks, they want to know who is the correct targeted customer for a loan. or who is the right person who would buy a credit card, pay the membership fee. So all of this data gets useful uh from their business use case as well. So this is the complete big data ecosystem. It is very important that you use the best of both. Having your human ability using your logic and prediction, the interpretation of a model is very important.
It is as in it is something that many times get ignored because you're using the computer applying advanced models to create results, but it's very important to not forget that human touch the human touch of not uh applying things Will be very bad because you will be actually missing out on the actual interpretation that only you as a human who has the correct domain knowledge will be able to interpret So using the best of the computer ability as well as your own ability to make decisions, predictions is very very crucial when working with data and data visualization. Three key things which just summarize exactly why data visualization matters.
The very first one is better decision making. If you have come to a conclusion, come to a decision, data visualization can help back that decision. And also, data visualization can pinpoint things which you might have overlooked while arriving to a particular decision. So, data visualization helps in better decision making. It also helps in a meaningful storytelling. Because you will be able to discover patterns, discover trends, you will be able to understand the flow of things over time, and that is very useful to identify trends. The last but not the least one is Data literacy. Just working about data and even if you're not too deep, if you are new to a particular domain, data visualization can actually help
uh improve your data literacy in the particular domain because you will while visualizing data you will see how data is placed taking a previous data taking a historical data you will be a able to understand how data worked at different points at different features across different features so data visualization helps a lot there as well While dealing with uh data, it often at times happens that there's a lot of unconscious bias which comes When handling data and actually data visualization helps uh omit the unconscious bias that we have So bias can be in many ways. If you are expecting that your company will do better
because you have invested in artificial intelligence, machine learning models and those are like the hype words nowadays. So it would just perform better. But that might not be the case. You have an inbuilt bias of uh using cool trendy words and technologies and associating it directly, mapping it directly to increase in sales. But that might not be the case. Data visualization can help you there. Also, it is very important that you be conscious about your unconscious bias So many a times most of us just assume that I'm somebody who thinks about all the perspectives, who thinks about all Points before actually coming to a conclusion, why will I have an unconscious bias?
But every one of us has some sort of an unconscious bias, and the best way to deal with it and to omit it is that we Should be aware of our unconscious bias. If we are conscious that we might have unconscious bias, we will give a second thought before arriving to the final decision. It is time that we tackle our unconscious bias. Consciously, once we consciously identify uh that we might have unconscious bias, the Only thing left for us to do is to avoid the bias at all costs and that is what will give you results which you would have never seen before Now, an interesting tool that I'd like to discuss with all of you today is Orange.
So Orange is an open source machine learning and data visualization tool. And this is actually very beginner-friendly as well as it has a lot of advanced features for somebody who's all already in the data visualization space So Orange has uh three key main features, just like the advantages of data visualization, that it it helps you identify hidden data patterns. It also provides intuition behind data analysis and procedure. And uh Orange uses since it uses data visualization, it helps it it acts like a bridge between data scientists and domain experts. So Orange is actually a tool which can either be imported like a library in your Python files
or be downloaded as a desktop application and you can import your data, work around with your data, which could be text images anything and then use like a drag and drop feature to actually play around with data This is very useful if you are completely new to data visualization, the field of data science, and you want to understand how data works, how various algorithms work in it. This is how a orange workflow looks like. Orange is actually a component-based data mining tool. The data analytics uh the data analysis that we do is done by attaching components uh which are called widgets with each other Here in this case you can see that there's a file component which will have all of the data that we have.
The data table shows it in a structured format. We are trying to use logistic regression, plot it as a scatter plot, all of this without any code and just connecting one point to the other. It is very interactive. It has uh the data can be interpreted very well. Considering this case where in the data is plotted as a scatter plot, if you want to see A particular set of data in your table and where exactly does it fall on the scatter plot? So highlighting those tables will highlight those particular data points in your data visualization Isn't it very useful? If you want to see how a particular state is performing, you will be able to do that, even though you're not using
states as a feature. So this can help you understand features, how it is being interpreted in any of the data visualization ways, and this makes it very very interactive. For example, if you are plotting a decision tree and you want to see the way it is represented in a scatter plot. In that case, color coding it will apply it to the plot as well. You can see that the classification tree, the red, the green, the blue is shown right in the scatter plot as well So what this helps in doing is that instead of using code and instead of actually diving deep into data, the structured the thousands of rows, you're actually interacting right there in the data visualization part.
So being it uh being very visual, it gives you faster results. You're able to test your hypothesis, see if the hypothesis you have made is actually being reflected in the data you're working on or not. And it has a very clear workflow design interface. Also, you do not have to worry if one particular thing can be an input to the other. Whenever there's a widget and you pull out a connecting A connecting point for another string only the ones that are compatible to the previous widget will be displayed and then you can search for it as you can see in the drop down here. What this ensures is that you are never attaching to non-compatible widgets
so you do not have to worry about anything like that. Orange does that for you. And it has great visualizations Any visualization that you can think of, be it heat maps, cell out plot, scatterplot, box plot, histograms, any and everything is right there So that makes it very useful for you, especially if you are a beginner or if you do not want to have even that piece of Python code again and again and write it uh differently for each of the things you're doing, you can have All of that being done through drag and drop in the orange tool itself. And working with all of data, you might have this question that since I've been talking about so many things that you can do with orange
But then there are so many number of choices and especially when data has a lot of features. So f finding the optimal feature pair to visualize data will be so difficult. How do I do that? Don't worry, intelligent visualization comes to the rescue here. Intelligent visualization is supported in orange. So for example uh through the score plots that are there in scatter plots, uh score plots help you find projections with the best class separation. I'll also be sharing an example of how we do this. But there are many things that you can play around in orange uh different things that you can try out which will help you in identifying the optimal k-value for example if you're working with k-means.
How do you do that? that all of that can be done right here in orange also uh orange supports reporting So to access the workflow history and everything, all of that is supported the orange. So you do not have to focus your time or your energy in keeping a track of how things have been going on Now taking an example of how exactly orange works, here is a snapshot of the orange canvas. I have tracked and dropped a file widget. Which has the retail and this is the ultimate thing that I've done. I'm applying key means and I'm plotting it in different ways through scatterplot, box plot and uh MDS and then saving it as a file. Yes you can save it as a file and then use it in your reports, in your presentations
or to share it with your colleagues so that you are on the same page of understanding information. So when we are using k -means and I am plotting it using MDS, so the first on the leftmost side you see the various widgets that are there in orange The middle pop-up is the k-means and the last one is the MDS plot. Let's focus on the second one that is k-means. In this case, what I have done is I have fixed the number the K value, that is the number of clusters, to be 2. That is why there are two clusters, one is red, one is blue. And it is pretty well separated. So because I knew with the data that I am working on that 2
is a great k value to separate my clusters. But What if there's some other data and you do not want to do the pre-mathematical calculation to identify the optimal k value? That is where intelligent visualization comes to your rescue. So uh instead of just trying it with k equal to 3 and seeing how it looks like, so The second point in K-means pop-up that you can see is a FROM field. So if you know that it could range from 2 to say 8 number of clusters, you will put in 2 and 8. And the cell out scores for each of these combinations will be displayed here
What if you are unaware of what a silhout score is? So if a data point is close to the center of a cluster, the silhout score will be very high because it belongs to that cluster. and as it goes far away from it the synop score reduces. If there are two clusters and a data point is between the two clusters, at that point the data point could belong to any of the clusters So the Silhout scores helps us give a value to the way To the number of clusters that we can have for our data. In this case, you can see that the sellout score for k equal to 3 is 0. 723, which is pretty high, which is actually the highest among all of them
So the optimum value of K in the K-means when we apply to this particular dataset is 3 And this is so useful because you do not have to do any pre-calculation in order to see which will be more useful to you You can directly use this functionality, use this intelligent visualization technique and find out the value of K. And also looking at it visually, you'll get a sense of uh whether it is the right thing or not for the data that you're working on the domain that the data belongs to you will have a more understanding of the data based on the domain that you're working in So that is how you can do great things in orange. Now an interesting use case of orange that I'd like to share with you is how orange is being used in space.
So the Hyampoosa 2 asteroid sample return mission uh wants to analyze the composition of asteroids that are there and the mission is using a number of things including orange in order to understand uh The composition of the asteroid. Let me show you how a small workflow, a small orange workflow, looks like in case of the asteroid composition sample. Yeah, it looks complex, but if you see uh if you look at it a little bit, you see that there are so many things that are happening and it is so easy because you are able to connect one widget to the other and being compatible you can directly connect it get the final output and that will make it um very
very easy Before I conclude, I'd like to say that since orange is being used in space and was developed here on Earth, why don't we use it to our advantage and make the most of data? Here are the references and attributes. So orange. biolab. si is the official website of Orange. You can download the desktop app, understand how you can import it in your Python functions. And this is a great book that you can refer to if you're new to data science and big data analytics. And with that, I'd just like to say that go ahead and make data speak volumes uh illuminate your data with visualization and make the most of the data the tremendous amount of data we have in the world today
And on that note, thank you for uh coming to my talk and it has been great sharing my knowledge with you. Thank you.
Data mining is the autonomous discovery of useful, novel, and understandable patterns in large datasets. Its four steps are retrieving data, preprocessing it, analyzing it for trends and relationships, and interpreting the results for action.
Discussed at 2:04Descriptive analytics explains what happened, diagnostic analytics explains why it happened, predictive analytics estimates what is likely to happen, and prescriptive analytics recommends what to do about it.
Discussed at 5:17Visualization presents data graphically, making trends, relationships, and patterns much faster to interpret than they would be in a table. It can also combine views from different sources and use color or other visual cues to highlight important information.
Discussed at 12:14A good visualization uses enough relevant, clean data, reflects the right story or domain, has a clear goal, and chooses a visual form suited to the problem. It should also highlight meaningful points and suppress irrelevant anomalies when appropriate.
Discussed at 17:40Orange is an open-source, beginner-friendly machine-learning and data-visualization tool. It can be used as a Python library or desktop application, where users connect drag-and-drop widgets to import data, analyze it, create visualizations, and test models without writing code.
Discussed at 26:25Orange can evaluate a range of possible cluster counts using silhouette scores. The value with the highest score is generally the best choice; in the example, k = 3 had the highest score, 0.723.
Discussed at 33:31Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026