Closing session
Published June 13, 2025
This video features Christian Tanul at DjangoCon Europe 2024 in Vigo, Spain.
Workshop: Functional LLM Chatbots - HTMX, Function Calling & LLama 3 by Christian Tanul
https://pretalx.evolutio.pt/djangocon-europe-2024/talk/ZGBQ9K/
Christian Tanul shows how to build a Django chatbot using Groq’s API, Llama 3, HTMX, Django Ninja, JinjaX, Tailwind CSS, and Pydantic-based Instructor. The application keeps conversation history in the session, sends it to the LLM for each completion, and uses HTMX partial responses and events to update the interface without JavaScript or page reloads. He extends the chatbot with server-synchronised dark-mode and full-screen controls, pizza-order CRUD operations, confirmation steps, typing indicators, and structured model output that can trigger client-side actions. He argues that simple prompting is often unreliable, so confirmations, structured schemas, reasoning fields, and careful event design are important for keeping an LLM’s actions controlled.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Alright, so I think we can begin. So first off, I won't waste too much time with my introduction. My name is Chris, I'm 22, and I'm getting getting married in two weeks. Now not saying this for YouTube Thank you. It wasn't for you to congratulate me, but uh in the past like six weeks, instead of helping my future spouse with the planning for the wedding, I was w working on this workshop like every day. So if this doesn't turn out well, I'm going to deny it for the rest of my wedding, uh for for the rest of my marriage, uh, for nothing. So If you manage to clone a repository, there's also this that you need to do. Groque is basically a team that has built their own hardware for running LLMs. And as far as I as far as I know, it's pretty much the fastest API you can currently use for running
Speaker 1: uh local models. I mean it's not local because you're in using an API but Uh they got their own hardware and it's it's free. Like currently you can't even pay even if you want it. Uh it's probably in beta or something, but it's really fast. And it is we're going to use Lama tree. You guys know Lama tree from Meta? Yeah. It's the 7 billion model is actually very very usable, I would say. If you know how to use it, it's much better than anything we've heard that we've had in the open source thus far. So you need to go to Group Cloud. Uh if you go inside the repository in the. env. example file, you will have the exact links to the is is the code visible? Should I make the font bigger? I think yes. But if you look inside of here in. enf that example,
Speaker 1: you'll find the links, the login and to the API keys. Let me try to make the phone bigger. Is this good enough or should I make it bigger? With the people in the back, do you see anything? Okay. All right. Thank you. So I haven't prepared like many slides. I don't want this to be like a talk. I want it to be a workshop. I did work a lot on the code. I spent a lot of time. Uh I've built basically it's three milestones, but you'll see like it's uh seven steps. Zero is just like the introduction to the project itself because
Speaker 1: there's some things that I want to cover. Um about the technologies we use. Now each step has a task and a solution. Task and a solution, task and solution. That's that's the whole repository. So we'll start with a start here. Actually let me show you what we're going to build first. Yeah, so this is the project, this is the final step. Now I'm doing this live, so I'm I'm betting that it should work. It's quite reliable. If it doesn't, I mean I've worked for six weeks on this, so hi. Uh what can you do? Okay, since we're talking dark mode, full screen mode, and pizza mode, okay. Uh can you toggle dark mode for
Speaker 1: me Okay, worked. Uh can you screen out? Uh we'll both up. Okay, so it's able to do like basic UI actions that I've integrated. It's not very practical of course, you wouldn't have to I mean who wouldn't press the fucking button and would ask the LLM to do it. But the idea is that uh Oopsie. I just want to showcase like this is possible. You can do things in the UI. You can for example maybe in the presentation I would ask you like make the phone big bigger or I don't know. There might be cases where you might want to change the UI to an LLM. I don't know This whole project, the whole idea that I had with it, was not to showcase like a usable product that I want you guys to use.
Speaker 1: It's like just the foundation of like you can do this, you can build on top of this, and maybe build something really cool with this It's all done with HTMX. This whole project has zero lines of code of JavaScript code. So all the interactivity that you see is just CSS, telling CSS. Um yeah uh and there's also the so this is about the client side stuff that you can do. But let's see um Kinetogle pizza mode. Then it goes into pizza mode. And in pizza mode we can actually place orders, pizza orders, update them, delete them or read them. So crowd functionality. So can you place a pizza order for me? It should ask me okay what I do want. We have this, this, this and these sizes.
Speaker 1: So I want a large pepperoni pizza Okay, now it asks for confirmation. I've added this step uh intentionally because it actually helps with keeping the LM a bit more tamed, otherwise it's just randomly trigger stuff and that's not cool. So I'll say yes. And it created it for me. And everything is down to HTMX. I'll explain later how. And now this is integrated like you can create uh pizza orders by yourself just manually. You can update them, you can delete them. And he should be aware of or it should be aware of uh how many pizza orders do I have. What else do I have? Oh
Speaker 1: what can you tell me about them? Okay, so it's uh it's capable of uh seeing uh can you make it medium size? Okay. So sometimes the API actually fails. It's usually because um it it's because of Grock. currently so it's not like a perfect solution yet but it's reliable enough. I mean do consider that this is a Lama tree based model. It's not fine-tuned So this is just through system prompting, giving like one example. Usually if you fine-tune it, let's say I gave it like fifty conversations, examples, or hundred conversations as examples, it's gonna be much more reliable. And it is possible with Lemetry you can fine-tune it
Speaker 1: So that's pretty much it, but let's also see how uh if it's possible to uh to can you uh make the pizza order 294 pouch Okay, it updated it. Now uh now delete it please. Okay, so all works thankfully Now let's go to the first start here. So the first uh so you can I've made two markdown files. So I made it so that even do not attend this workshop with people that don't even look at the video about this workshop should be able to go through this workshop. There are two markdown files. There's the README, which is it doesn't change through branches, it just showcases like this is the application.
Speaker 1: Uh these are the weight requirements. So by the way, these are the requirements. You should have git, I suppose you do. Uh does anybody have Docker desktop? I should have probably said faster. Uh do you guys have Docker? Yes? Is there anyone that doesn't have Docker? Okay. I think you should be able to run it anyway, but you might have to run pp install requirements and then python manage. py run server. Okay, so uh and then the cloud account. Now like I said it is useful if you guys know the specifical. If you know some HTMX it's useful. If you've used the OpenAIS API before It's very s very very similar. So most LLM APIs work in pretty much the same way. You have a client that you create completions from.
Speaker 1: So I'm showing you the more complex code right now, but I just want to show you the part where I generate the response. So here it they all have like a chat. completions. create You choose a model and you send the message history. So with each new message you basically send the whole previous history so that it knows what you just talked to him So basically what we need is just a list of messages that we keep updating and we keep sending it back. Does it make sense? Alright. So um so clicking inside the markdown file. Yeah, so that's what it is. Now you've already cloned the repository. Have you guys managed to create a Grock cloud account? And created an ATI key? Did anybody not
Speaker 1: smell? Okay. Oh, it's internet. I see. Well keep trying and and meanwhile I'll continue with explaining what we're going to use. And now this is a uh I'm curious, would you guys put my head on a spike for using Django Ninja with HTML? Because Django Ninja how many of you go guys know about Django Ninja Okay. I personally find it like a really cool project. I use FastFast API and I liked it, but I don't like not having the Django batteries. And having like both is such a cool thing So Django Ninja was initially created it is it is supposed to work with JSON. It uses Pydentic for validation and stuff.
Speaker 1: But I used it just because it's less verbose and I wanted in this workshop to have as little plotter in the code as possible because I want to showcase so many things. So Yeah, I hope you don't hate me. Now HTMX of course. Um Okay, so there there have been many talks about HTMX. Uh So those of you that don't know about HTMX, do you not not know anything about HTMX? Should I try to explain it in like one minute? I think I can Yes? Okay. So basically normally without HTMX, let's say we say we have a typical Django project and we use a Django templating engine, right? And think about a form. We have a form, right? It has two fields, username and email. Normally, without HTML, you would fill up the form, you press the submit button, then
Speaker 1: the page reloads, it does the post request, and we get back the response. Either the form is good or it's not. Now Now with HTMX, things are a bit different, but not very different. You have some special attributes. HTMX means HTML extended So you have some special attributes, one of them is hx post. So you would replace the form 's normal action attribute, which has a URL, to hx post, and you choose your URL. Now, when you submit the form that has this HTMX attribute, instead of doing a page refresh, it does an AJX request, so it doesn't change the page. It does an Ajax request, it handles the server-side code about like the form submission.
Speaker 1: And then it sends back an HTML partial, so just a piece of HTML that gets injected in the existing page. So you don't have to refresh the page. So this gives you like time is interactivity like a dynamic web app um without a lot of overhead. Like it's really very simple. Once you see it it's that that's why most of the people that learn about HTMX usually join the cult as well. Because it's so simple to do really cool stuff. And I personally I started working with HTMX about three years ago. And it was very like it was very underground even then, but now it's like I I would say it's at the top of the underground because you can hear many talks about it right here Uh but I'm glad because now we have much many more examples of it.
Speaker 1: So YHTMX, complexity bad. This is this is actually from a blog from the creator of the HTMX which is really funny. And it actually actually does provide some really wise advice from a senior developer. So I recommend it you read it. Now we will use GinjaX. This is another cool project that I found. It's not very known. You don't have to use Ginja X. I just wanted to showcase like this is possible Uh there was this talk yesterday about uh I don't remember her name. About Alpine JS and HTMX and uh how it works in production, right? And one of the issues was that your templates get very cluttered. You have too many attributes. You have like HTMX attributes, you have L. js attributes, just too much behavior in the templates. It's difficult to see what's going on.
Speaker 1: Uh okay, uh I want to ask a question. How then answer honestly, how many of you love working in Django Okay. How many of you love writing Django templates? Okay. So this is not uh not a critique at uh Django of course. I'm just saying like it's at some point Especially when you work with HTMX, you need some kind of composition. And that's when many projects, many cool projects came, such as maybe you've guys heard about Django Components uh snippers and Django template partials, which is from Carlton. Is Carlton here? I'm not sure. Okay. It's a cool project. And it actually simplifies this process of not having to uh create uh so much uh boilerpl um so much template code.
Speaker 1: Okay, so Ginja X, simply stated, is just Ginja 2 with syntax sugar for macros. That's it So this is an example from his the developer space, the maintainer's page. Oh. Yeah, does this go off sudden or like uh frequently? Sorry. Okay. Is this from the HDMI? Sorry guys. Okay. If you guys look at the readme, you also have it on your on your laptops, you will see this uh if we just can't do it with this way.
Speaker 1: So that's that's the the b side to side example. Now I picked I I looked for this project because I I worked with Django for about four or five years and I've seen most of the cool parts about it. Like I love writing the clean Python code in the back end, but the Django template is always like a I didn't like in writing them. And uh I used Ginger 2 as well and I personally I liked it a bit more. Um But then I tried Astro. I don't know how many of you know about Astro. It's like a JavaScript framework. It's not like a front-end framework, it's like a meta framework. You can actually it's pretty I would say it's a pretty cool skill to have in your pocket as a Django developer because some projects you don't need to uh make a whole whole backend for it also. Yeah, but in Astro they use JSX, this this syntax where you have composition
Speaker 1: Yeah yeah please do. Well we have uh composition by creating these uh tag like structures That you add in your templates. And now some people might not like it. Like there are arguments against it. I would say uh firstly when I look at a ginger template, I think it's good that when you see curly brackets, you know that's ginger And when you see template code is like HTML. So in this project that's a bit blurry. But I personally enjoy it more because it's easier to read. So why do I choose GingerX? Because with HTMX you can do this cool stuff. You can abstract the c the HTML parts that You don't need to see at a glance when you look at your page. So for example, this thing this is the index. html, the main
Speaker 1: the top level file of the template. You can see the structure very clearly and you can see the behavior very clearly If I wanted to work more on what's under the container, what's under the chat message, I can c I can just go to the component that includes that HTML But on my top level file I have like this simplicity in the structure and I can see the behavior clearly. So that's why I like it. I'm not pushing it on anyone, I just want to showcase like this this project exists. And if you guys like it, try it. Okay, I made a blog about this as well. So if you guys like it, uh you can read it. Uh ECSS, I suppose you guys know about T CSS. I don't expect everybody to use it But uh it's pretty nice because it pairs well with the locality of behavior that HTMX provides.
Speaker 1: Uh and you also can do like pretty cool stuff with uh advanced stable CSS. You can b build variants. Variants are this like cover dark uh prefixes to classes and I build this HTMS request on. I give a CSS selector then I say block. So this is like the typing indicator that we will see later, which only shows when a HTMS request is being uh in flight. Okay, and then the script loud. And I think that's it. Oh actually no. But there's the instructor, which I will actually explain later when we get to that part, because it doesn't make sense I should explain right now. So let's go to the one integrate integrator LM task. Um short.
Speaker 1: Only already twenty minutes passed, sorry. I wanted this to be much shorter, but you had to have some context about the project. So what do you see now? Um Yes, sure. Um think for telling me. Did I open up the camera? Okay, so what you will see once you go to the first task is a chat interface that you can talk to, but nothing answers. So I've implemented the chat itself with HTMX. You can send messages. All you need to do is work on the view. py and the index. ginger files.
Speaker 1: Now you will see uh pretty irregular template structures. Uh again, don't put my head on a spike. I picked this up from Astro and I personally with this Gingerx I found it Easy to understand. But also basically includes like base. html. Components is like components in which I don't put behavior, it's just a component. I add the behavior with the extra attributes. And uh if you want to look at the pages, you know the pages and index. And see here you'll see the two things.
Speaker 1: Uh is the chat container, which is empty. Uh it has a placeholder and it uses selling CSS to display it only when uh actually this doesn't even need to be here because we don't have the typing indicator yet. Basically it's only sh uh it 's hidden once we add chat messages. And this is how we add chat messages. We have this form. It does an HX pose to add user message. It targets the chat container above and then it uh adds the response, which will be an HTML, at the end of the chat container. So let's see the view and what it looks like Okay. So what we do first, we get the we save the chat messages in the session, the session state, the server session. Okay. It's in the cache right now. I haven't added a database for it. It's just in the session, in the cache session.
Speaker 1: We add this is the expected structure that most LLM APIs expect. It's two dictionary with two keys. One is the role and the other is the content, with the message content. Here we have the role of a user because we are adding a message as a user and we add the content which is from the form because uh you will see that here's a chat input with the net it's basically just an input under the hood with the name of message. So it works like in typical form. It sends the name with its value which we add as we type here. When I press enter It goes to the view, it's added to the chat history, it uh it marks that the session was modified and then it uh renders the chat message component Now I've created so
Speaker 1: Ginja X has its own way of rendering templates and I've built a wrapper like a function that more resembles our typical Django way Now I couldn't get be around the fact that it doesn't expect the whole path and it doesn't expect the extension. I would have preferred to be able to write pages slash n index. html I'm actually talking with the maintainer right now and I want to contribute to the project and see maybe we can find a better way to integrate it in Django. So this is this project is not for Django. We built it for any other like it's agnostic, you can use it with Flask, you can use it with whatever. But uh this render utility function is I try to make it similar to Django's, but it also includes the headers, which you we will need later
Speaker 1: Uh okay, so for the people that don't understand HTMX, does this make any sense? Like the the how the actual uh adding of the chat message works? It has this form , send a message, it does the pulse request, uh asynchronous uh through Ajax. And then the view just renders the HTML and it adds it inside the chat container before the end. And then we have this modifier which basically says But also scroll the container to the bottom. It's very readable, right? Because otherwise we would get messages and they would go under the screen and we wouldn't see them So this is basically the chat. So I wanted to give it as an example because you are now uh going to try and implement the assistant part.
Speaker 1: I've created I try to make it as clear as possible which each step you have to initialize your client, get the chat messages just like uh above, create a system message in a specific uh format Which I've specified, then create the completion, extract the message from the completion, and then append it to the session. And then you need to render the HTML just like above Uh do you guys think you can magic manage to do this? Like the people that have it actually the people that haven't worked with HTMX, I'm just going to come and help you, no problem. Uh but it would be great if many like more people that don't know HTMX group so I can explain to all of them and the rest like for I'm sure that for some of you this might be like super easy task, but for others not. So I don't know how to it's my first workshop. I don't know how I should proceed.
Speaker 1: So uh what do you guys think? If you want, I can just like leave you be for like I don't know fifteen twenty minutes and see what you can do. What do you guys think? Well give us a couple of minutes and we'll try. Alright, sure. Okay. Well uh then in the meantime I'll just put this here. Actually, wait, no, not this one. This one. So this is the first task. Now which of you would you like for me to come and try to help you?
Speaker 1: I have no problem. I would be glad to help if I can. Not just sit here and look at you. Okay If you guys uh it feels difficult for you to manage this task. So let me clarify, make it simpler. Um at the end, when you render the chat message of the LM You also need to add a special attribute uh in the header. I've added those two second actually Oh yes, no here in the users uh in the user's view. You need to include an Ajax trigger response header. This is a pattern from HTMX That is pretty cool. That lets you also send some events that the the client can listen to and trigger other HTMX requests from it.
Speaker 1: So here if you look at the um Yeah, so this is the this is basically the pattern. From now it doesn't showcase how to do it from the back end because this is like a backend agnostic in H in Django Well we just need to set the headers. So for example this it will be HX trigger at the system message. Now in order for the front end of the template to listen to this We would need that inside index. jinja here where it's uh trigger resistant uh DAD you 'll have a JX trigger Add the synth assistant message from body. So when this basically tells this div that whatever logic, whatever HTMX logic it has
Speaker 1: will be triggered when it receives this client event And so here we would add hx post at assistant message, hx target chat target chat container, just like above, so the same thing, hx swap before end And before end means b so yeah, go ahead. Yeah, so the thing is the basically the server sends uh an event in the headers And if I understand it correctly, HTMX picks that up and adds that event on the body element of the page. So you don't have to listen just from from body. It's not simpler, but you could achieve this without this special HX
Speaker 1: trigger HX trigger. I could say from user message from So the form that's the sorry form the form that's just above here. Now receive the HTMX after request. This is a special uh special um event that you can listen to from HTMX. HTMX has many events that you can use. They're very useful. So basically what I say here is after this request receives a response, so after we add the user message and it also updates our server sessions chat history. You can generate the assistance message. So if I've done this correctly, let me see if this will work. Actually no, because I also have to implement the view itself. But it should work. So but I also I wanted
Speaker 1: to show you that I think it's easy easier to just see that you send a custom add a system message header from the user view, which will trigger the system to respond from the front end. Pretty much, I think conceptually kind of yes. It's more like when you add a user message, it's just a trigger. Like now you can go on and do the other thing.
Speaker 2: This is what someone was saying in too late to see it's linked to this and that's linked to that and how smart that's not
Speaker 1: No, yeah, actually they are no sorry, the buttons. I'm I'm not sure to the answer to be uh to be honest. Like yes.
Speaker 2: Yes.
Speaker 1: So uh um message body? Oh in what way? So what do you mean? You
Speaker 2: receive a header from from the response trigger.
Speaker 1: No, we're just calling the we're just calling the request. But if the trigger just says, okay, you can go on with the request. That's it. So actually b so we that we don't waste too much time because I think we only have like oh no actually we have fifty more minutes. All right I'm going to show the should I show the solution? Or should we just I just switch the branch with the solution? I'm sure like some of you didn't even manage to clone the repository. Did you? Or did you guys manage to get here? Well uh the thing is, I'm sorry this happened now, but I tried to make this whole repository so that even if I fainted during the presentation you could go on, meaning i it's all written down.
Speaker 1: You will have the progress. md for each step, so you will see what happened, the context. This basically just shows a task. But when we go to integrate solution, if we check out here, to make sure this doesn't Upset it. Okay, so now you would see a progress that MD that says okay this is what you should see now. The system should respond. And I've also added this uh typing hint uh typing indicator so uh so throughout the steps I've added some simple some additional cool things that you can look at if you want. I didn't want to make their main part of the presentation because they're not. You want to see like you can do this. Like you can add a typing hint in the typing indicator while he's uh generating a response.
Speaker 1: Okay, so let me show you the solution now So let's start with the view. We initialize the girl client, right? We add the API key from the OS environment Then we get the chat messages from the session just like above. I've created a random system prompt saying to just create a high queue revs in the user for changing the system prompt. And then we try to get a LN completion. So here you'll see that I'm adding a list with a system message, which is this this specific format that I've told you, like a dictionary with a row that's system and the content being the prompt and then the chat messages which is basically also just a list of such formats with user assistant user assistant Now after this we get the completions
Speaker 1: content because you need to go into it. There's more details that you get from the API from the LLM. Then we update the session. And after this we just add the chat message of the assistant and the context is role assistant. And this is underscore content because if you look at the chat message that Ginja component here, it has this uh content uh variable and if you want to inject it from Python you have to do on double underscore that is the way Gnjax chose the guy chose to do it Okay, so um and on the front side uh this is the trigger assistant solution What what this does basically it says HX trigger
Speaker 1: when you get add assistant message from the body and it's on the body because that's where HTMS basically always throws those server events, it sends them to the body. When you see this on the body, do an HXPOS to add assistant message and target the control container to where you are going to put the response And the swap strategy, that's what HX swap just says. What is the strategy? Like how do you need do you want to replace the whole chat container? Do you want to edit at the end? Do you want to edit at the beginning? So it gives you flexibility in that way And before end means before it's end. So we have the chat container and we want to add the chat message at the end. And then it says scroll the chat container to the bottom. That's what all it does. Does this make sense Okay, great.
Speaker 1: Now if you guys manage to go to the third milestone, I mean If you finished the first one or if you if you still want to keep trying with the first one no problem I don't you don't need to keep this pace but if maybe some of you finished already you can go to the next one where I've added some new stuff Now how many of you will remain at the previous step? Because if many of you, I won't continue explaining some more stuff. I will let you just focus Uh so how many of you want to switch to the third branch now? Okay. Okay, then I'll explain really quickly uh what is changed here. I've added a little bit of stuff
Speaker 1: What you will see is that you now have two chat checkboxes up here. The first one basically turns on dark mode, the second one turns on full screen mode Everything is implemented with CSS. All it does is CSS just looks for whether these these are just some chat checkbox simple inputs. Let me show it to you. uh here. So this is a dark mode checkbox. It has a name of is dark mode. It's a checkbox. I turned off autocomplete because it's it's made some weird behavior. And I just said if is art mode settings checked. And what it does is It works even if you delete this, but then it doesn't uh sync with the server set state. So why do I want to keep the dark mode server set
Speaker 1: server state uh in the session state in the session storage. Because I want the assistant to be aware of what the dark mode is and what the full screen mode is so that when I'm talk talking to him, I for example I have dark mode on and I say turn on dark mode I don't want it to say okay enabling dark mode and then turning off dark mode. I want it to say okay dark mode is already on, I don't have to turn it on. Right. So that's why I'm I've I've basically synced it with the view, with the session in server. by using HTMX. And this is a slightly different pattern in HTMX. I've I've added this new attribute called HR Select. If there's a cool library from Carlton Gibson called Django Tanti Partials, but unfortunately
Speaker 1: I don't have this in Gingex right now. I plan to contribute and add something like that because I'll explain to you right now why. So these are the views for the toggle dark mode and toggle full free mode. As you can see it's not Doing too much. It just has two keys inside the session, is dark mode and is full screen mode. All it does is it swatches. Sweet. So it's a Boolean. If it's true, it makes it false. If it's false, it makes it true. That's it. And then it renders the entire index page, but only passing the e start mode uh context. Now if you look in the index. jnj, you'll see some weird things here appear In GinJas you have to define what props you are going to pass to a specific template. Now we have a request which is expected by default, but I made sure that it is passed from here.
Speaker 1: Uh we have some chat messages, but they also have some default values. So don't be like don't feel like this is so weird because we are rendering the whole template and wasting a lot of logic that we don't need to do because If we only pass the dark mode, false, it will only change the dark mode, and the rest will be as if it was a template that wasn't uh just plain HTML, no logic, because it will just use the default values There will be a workaround for this so that we would be able to do something like index dark mode partial or something. And then we wouldn't send so what this does is If we look in the network tab of the console, it will send the actual entire index page. But we will use HX Select to only select the actual input
Speaker 1: From that response and only that gets placed in the new in our current page. Now, this is not like a huge performance issue, though it seems weird, but it's not a big deal, to be honest But I personally don't like it. I just want to do the you can look up Django Template partials, it's a much much more elegant way. Also it would make it easier when you look at the view Like in this way you say why does it render the index? Like only when you look at also at the code that says oh okay it selects only the input, only then it makes sense. So that's why I don't like this approach. But I'm glad that at least we can do this because otherwise I couldn't have even done this pattern without HX Select. So that did I explain correctly. Like this is this is how dark mode and the full-screen mode toggles
Speaker 1: work When you press on them, it renders the index page with the new new value and it's it gets added to the page. And the CSS on the page is just looking for, I'll actually show it to you real quick This is like a tailing code. These are the variants, but you can see that it's quite readable. It says Apply the fullscreen classes, there are some full screen classes in the code, only when the HTML has an input with the name of each full screen mode that is checked. So that's all it is in David CSS. And then if I add any type of classes that have that specific prefix, they will only be included when I press this or when I press this. So that's it once you get a hang of it is pretty straightforward.
Speaker 1: Yeah, so there's our clockwise and in this task what I want to have you guys try and do is Use the Python instructor. And I'm going to explain really quick what the instructor does. Essentially, in a single sentence, it lets you get structured output from the LLM So LMs normally just respond with a with words, right? You say hi, it says Okay, I told it to like scream that it wants to be able to do dark mode and full screen mode. It was just for fun But it only responds with text. But what I would like is for it to say something like um message
Speaker 1: and whatever message then uh I don't know maybe is dark mode true or I don't know I just want a structured way to get okay what does it want to do Because if I tell and then I would prompt it in the system prompt that when you see that the user specifies he would like to change the dark mode Add this is dark mode true in the response and then I can parse that, turn it into my logic that actually changes the input with HTMX, and uh it would change the actual client. And There are not like super many steps, although this might sound a bit uh like I can't understand things when somebody somebody explains them to me speaking. But when you look at the code, I'm pretty sure if you look at it for like fifteen minutes you'd understand what's going on.
Speaker 1: So this this Python instructor Library does this for us. It does it does the whole parsing. We just use Pydentic models. And I think that's really cool. Because right here I would be able to just say message string. And then after I just a second. So we have the Gro client already initialized and all we need to do to apply instructor to it is to do instructor. From Glock we pass the actual client and we say mode instructor json and from now on All it does it it lets us add an extra attribute here called response model
Speaker 1: or then response And from now on, this variable, whenever we get it, it will be an actual instance of our Pydentic model with whatever structure we gave it. Like if I would say, look, toggle mode, toggle. mode Boolean I could do LLM toggle dark mode and it would give me either false or true depending on what the LLM answered. So this is really powerful because you don't have to stress about parsing all that JSON and making sure it's in the right format and all that. And it also enables you to do Now the way that I expected, and I think uh it's a good way to do this uh task right now would be to create a client events list.
Speaker 1: That would expect some literal meaning exactly two variables because we don't want it to think that it can toggle anything. We just want it to know that it can toggle dark mode and full screen mode. So we say literal , say toggle dot mode, and we'll go full screen mode And then it will be as simple as just doing um adders equal hx trigger LLM response client events. Actually I would need to do uh. join because the HX trigger expects a specific format, you can look on the reference. But that's all I need to do, and then suddenly My LM can pass HTMX client events directly to the response. Which is pretty cool.
Speaker 1: I mean you only need to add two lines of code and suddenly it can do whatever. Like if you had this HTMX already implemented in your your LM can actually do it now. So okay, I kind of spoiled the solution to you, but you can still uh go ahead and try to do this yourself. And if you want some more uh uh so if if it's not clear how the toggles work and all that, inside progress. md I actually created a visualization. You can see exactly the user presses the checkbox, it does the pulse request, it toggles it in the session, it renders the template with the new state
Speaker 1: And then it updates only the input with HX Select. And then the same thing with the full screen. So maybe it's simpler for you to understand the conceptually by seeing it. So you have all this inside the progress. md Should I come to anyone for help? Prompt. Don't mind. If you want it you can actually contribute to the repository with that solution showing like you can also do locally. Like this will be Open source, anybody can use this project freely, you don't have to do any credit, you can use it in your projects, I will be happy if you do. I mean if you just tell me like look I use this for this, I will be happy. That's it
Speaker 1: By the way, if some of you already implemented it, so maybe you're more advanced with this, you can try to do the challenges. I've added some challenges and uh progress. md but I've also added him in line. And they are you can try to include a reasoning step inside the Pydentic model. This will make the LLM when it gives your the response it will actually uh reason before it gives an answer and from my experience trying it uh it makes the responses much more reliable. Just having a s a simple internal process like uh the user wants to turn on dark mode, I should turn on dark mode. That's it and then it does it. Because otherwise if uh for example previously without a reasoning step, what I've seen is you could
Speaker 1: we would say hi And it would total dark mode. It's like a i if uh explained in the next uh progress. md. Actually let me show you real quick. So , Thank you. So it has a very impulsive behavior. Like if you if you tell it that it can do X and you don't try to tame it somehow at the same time in the prompt, it will just trigger it
Speaker 1: Constantly. They would just hi. You say hi and then dark mode, full stream mode. How are you doing? Uh light mode, small mode, so it just keeps doing it repeatedly. And this is an example. And you would just turn them on constantly So you can solve this with some system prompt uh by changing the system prompt, but also by adding this reasoning step, uh which actually like in this example you can see it uh for debugging debugging purposes I added a whole JSON You can say the user requested to toggle blow and then it does it. So it's a pretty cool thing. You can read some of these after you finish the solution. I've I've actually mentioned some more stuff It's basically the issues that I've hit while doing this and the solution that I've found.
Speaker 1: Technically technically it should work, but uh I'm pretty sure the most data that it received was in English. Probably Spanish is also alright because there's a lot of data in Spanish. But for Uh sorry?
Speaker 2: Three definite levels like for
Speaker 1: Yes it does. But it's I would say it's not as good at instruction following in other languages. Because for example, I've built this project And I'm from Romania and I build this project. It's called intrapolegia. ro. It means ask the law. So and in this is in Romanian. And I've tried I actually tried uh making um a system prompt in English and then making it respond in Romanian although the instructions themselves were in English um and I tried to make the instructions in Romanian And this is Claude, Cloud 3 uh model. It's pretty good, actually very good. I use Cloud3 Haiku, which is the smallest one, it's like one fifth of the price of GPT 3. 5. And still it gives very good answers. And uh From my experience at least
Speaker 1: with Claude, it was good enough. I mean I didn't see a big change and considering I wanted to kind of guide it towards the tokens that are specific to Romania because it's the Romanian law uh I giving the instructions in English would have done a worse job even if it even if it understood them better. Um I don't know if that makes sense. So uh the thing is with each each token, each word you give it to it, it kind of guides it in a specific area in its knowledge because it The LMs are like a an ocean that is full of knowledge and you just need to do the right combination with the words to to touch gold. And that's basically prompt engineering. Yeah. Tabolia is free by the way. It's I made it to make profit. I don't think you made it.
Speaker 1: So I'll explain this to everybody because maybe it's not maybe it isn't making very clear, so sorry for that. Um I mean if you do get to this point, if you manage to make it to turn on dark mode and full screen mode, but it it isn't aware of sorry It isn't aware of the current state. You need to find a way to pass it the context on each um on each generation of this its answer. And you can do this by just changing the system prompt. So the system prompt doesn't need to be the same each time you send a message. You can actually build it up dynamically, meaning in the first message it might know that dark mode is off and full screen mode is off But for the next message, I might change the system prompt when I send the whole list of messages and it will know that the dark mode is on.
Speaker 1: And then it knows that it we can turn it off. So what I mean is uh inside the add assistant message. I think at this point I think I've said uh this is the solution already But uh I think it's okay to show you this. So I've moved m I've moved I've moved the prompts inside a separate prompts. py file and inside of here I have the context and rules And then I have the reasoning instructions for the reasoning step, which really helps being uh getting a more reliable answer. Then I created a function. It's called create client event instructions. And it gets two very uh two parameters. Is dark mode and is full screen mode. And basically whenever I'm creating a system prompt, I'm passing the current session state
Speaker 1: and then it knows this is the current state, dark mode, false or true. false or true and then I also this also seemed to help it better is saying this is what you can do you can add double dark mode and then you will change dark mode to not is dark mode so it's very clear for it This m actually makes a very big difference. Like it needs to you need to be as clear as possible, otherwise it will interpret you in different ways. So if you add this as well like change dark mode to not exact mode, uh it will be quite um consistent with the answers And then what I do here is when I create the system message in the content, I just add context and rules plus reasoning instruction plus the result of this function in which I add is dark mode and is full screen mode, which I just got here from the session state.
Speaker 1: Does it make sense? Yes.
Speaker 2: See you put a full screen and a lot more than so
Speaker 1: Yes. This would become an issue. Yes, because you pretty much fill up the context. And that would be even more of an issue in the next step. But we will have pizza orders and we want it to know all the pizza orders. So if you have 300 pizza orders That's a solution for another day. Right now I want you guys to see like I've built this starting point. You can actually use a retrieval augmented generation for that You can sell it you can actually have so let's say I had this um super agent and then I have some other smaller agents. One of them gets the pizza orders and when he figures out that he needs to know what the pizza order is He makes another prompt he tells the agent to take them, then he gives the answer. So this is possible.
Speaker 1: And I'm going to give you another hint. If you try to add the reasoning step, it is very important where you place it. Like the way you structure this response. It will actually generate the JSON in the way that he added the field. So add reasoning and then the client events and then the message And this is the way you want it to be because if you put the clients first and then the message and then the reasoning, it will basically start by saying total dark mode, then the message will be according to the action that it took, saying I just followed our mode, and there is reasoning will be like a post-action justification instead of an actual reasoning. Like you want the reasoning to be at the beginning. So it will reason like what is going on, what I should do, and then it it will do the things. And I've explained this in the progress.
Speaker 1: md. If you look at the progress. md in the solutions, so after we finish You will see this explained very clearly. But uh the order of the fields is very important because uh basically each new cotoken depends on the previous ones. So it kind of directs it in a specific direction.
Speaker 2: Two instructors to get the JSON It just creates it.
Speaker 1: So basically what what pi what the instructor does under the hood is it takes the identity model that you give to it. So when I add this response model, LM response, it wraps my system prompt in another as little as possible prompt saying you are expected to respond in this exact format. like the guy that made a library try to minimize the amount of prompting but it's pretty good and you can actually actually also add some max retries So it if it doesn't provide a good piedenic uh answer, it will try again. And you can say like try ten times, you know So that's what happens under the hood. That's how it's possible to create an instance of the Pydentic model. Because the LM responds with the JSON, it's then parsed into the Pydentic uh instance, and then you can
Speaker 1: go ahead and use it
Speaker 2: Mike the chance of
Speaker 1: Uh no, so basically the instructor just says that reasoning is the first thing it's supposed to do, but when it actually gives the answer He starts creating the reasoning itself and that's the part important part. Like the the con the content that it includes in the reasoning is what's important, you know.
Speaker 2: Yeah. I think turn on.
Speaker 1: Oh yes Uh yes, of course. So if you look at this, you can do lm response message or you can do lm response. json because it's a Pydentic model, you can just do that JSON. You will see the JSON format of the And actually I recommend you do this because it will be easier for you to see what's going on. I have a question, maybe nobody did, but did any of you hit a roadblock where the LLM works for like two or three messages and then it stops working? Okay, if not, no problem. I was fine. I hit that roadblock and it's it took me so long to understand why, but it's really funny.
Speaker 1: I think we only have like fifteen minutes left, so yeah. So let me see if
Speaker 1: see if it's an issue with the actual branch because maybe it's the same to me. No, for me it me it seems to work, but there it Oh yes, yes, yes, that's true. Yes, let's let's see real quick what Docker does else because um so I copy the requirement add requirements, I install them, and that's it. I don't know, honestly. You can I mean you can't install Docker with this internet, but uh I recommend it it's much easier to run projects with Docker
Speaker 1: You probably like it probably does the right thing. So the issue is not it from your part, it's just the CSS So should I uh go real quick to the last even if we didn't finish just so I can cover everything and you guys can see it, like the crowd functionality, the pizza mode and all that? Okay. So Oh since we only have like twenty twelve minutes left. Is it twelve minutes left? I mean that yeah, yeah, okay. Um I'll just go to the solution and show you the solution itself. And you can if you want you can go do it uh the task at home or whenever you want. So And also make sure to read the progress from the fourth task because I've added some very interesting
Speaker 1: stuff about uh the the order of the fields and all that. And uh Race conditions that happen when you you need to add a delay to the toggles, but whatever. Okay, so I think I didn't change the this title up here, but that's matter. Okay, so if you want you can look at this and see what the final result would be. So we have the Gro client, we have the instructor. Then we get all of the client state, all of the sorry the session state.
Speaker 1: Now in this in this branch I also add a pizza orders app, which actually has its own model It has the views, the crop views, and it has this services. py file, which basically includes what you would normally put inside the views But there's a reason I've moved it in the reservices. py. It's so that the LLM can actually reuse it. I don't want to pass the actual view functions to the LLM. I want it to just do the the the crowd functionality. So this should work And let's see uh what happens when I tell it to do something. Let's say I ask it uh create a pizza uh large oop Create a large
Speaker 1: peronic pizza. I say yes. Okay. Okay, it doesn't work. Okay, whatever. So what it would normally do is it gives the pizza orders from the database as the values just like just what's important, the ID, the name and the size. Then we create a system prompt with this is the a similar function to the one from the previous step, generate contextual information. Here I pass what is important for him for it to know if it's dark mode, if it's full screen, if it's pizza mode, and the pizza orders, so that it knows about the pizza orders. And then here I give it the current state of each thing and what it can do and what it will do. And then there's this new step of pizza orders
Speaker 1: server. It's basically the current state will just print out the list of pizza orders. And then it said these are the possible actions you can take. And uh Then inside of use par there's not a lot of code. It could be abstracted in its own it like if you have more of that functionality, you have multiple models, you might not want to keep everything in the same view. But I thought for the sake of this workshop it's better to have everything explicit so you can just look through it easier. So uh the way that I've changed the Pydentic model is so I have the reasoning, right? I have the server functions, which is a list uh of tuples. And why a list of tuples and not of not just a dictionary? It's because in a dictionary I would have to say okay the keys can be create, update, delete
Speaker 1: pizza order, and the answer can be pizza order I pizza order in Tople topple int pizza order in or int. And then basically if it's not very clear to him, it could call create pizza order and just give me an int. And I don't want that. I want clear that this is exactly what I want. cre create pizza order with the pizza order ID payload, which is just another pie pay uh it's it's another pyentec model. Um So uh that does make sense why I chose this this way. It's a list of potatoes. Now after we generate the response I have this part where I say for functioning and data that basically payload of the function like what it includes in LM response server functions Match the function.
Speaker 1: If it's createpits order, it calls the createpitsolder function. That's from the services. py that I just told you And I say the payload will be the data. If it's an update, it's order ID and Pizza Order because this is the expected payload. Then if it's delete this older, I just delete this order and I give the order ID. So basically I'm just letting I'm literally calling the function that it just told me to call with the payload that it told me to use And then um now I've Don't know if I get the time to explain this, but uh I've also used a slightly different pattern of HTMX for the pizza orders. So here in the chat, when you send a message,
Speaker 1: It adds the message, uh it creates the chat message HTML and adds it to this existing container. But there's also another approach. where you have um an extra view. So this chat container would include an HX. Actually let me see, I think I have a visualization visualization for that as well. But I don't know if it's here. But basically the chat container itself would be hooked to a view that would get all the chat messages It would expect a trigger of chat messages changed or chat messages updated. And whenever I send this trigger from the server, it will update it with the with the visual with the exact representation of the what the server actually contains.
Speaker 1: And I believe although it's a bit less straightforward to understand, it is uh Safer way because if anything happens on the server with the first solution, so let's say for example, right now we're using the cache for the chat messages, maybe the cache just resets. We suddenly In the back end our chat history has been removed and we give a new answer to an empty list. But on the front end we see it as if we have this whole conversation and suddenly it just doesn't understand the conversation. Like why does this happen? So we can we can see exactly what's going on in the back end. But with this other approach, whenever I send a message, it sends it to add user message. The response is an empty response, no HTML, it's just the trigger, update the chat. And then the chat updates itself with the
Speaker 1: with the actual state of the chat messages. And then when the assistant responds, uh I think it's just the same thing. It just adds it and then sends the chat message updated and then the chat itself updates itself. So this is a Common pattern in HTMX, and you will see it if you look at HX trigger in the reference in the documentation of HTMX. So this is what happens in the pizza in the pizza container. And for reason why The first way is also more difficult is because for example here we might want to add one, okay? So it's okay, we just We just add the HTML of a new pizza order in the container. But if we want to delete it, what do we do? You can do something, you can actually add some other headers that
Speaker 1: tell it to delete this ID, but it Suddenly the code will become a bit more messy. And I think the first approach the other approach of having the extra get chat messages or get pitch orders is more straightforward and it's safer So yeah, if you will look to the project you will see all the patterns. I've tried my best to like include as many HTMX patterns throughout the code so you can see what's possible. Because for me the way I've learned HTMX was just seeing other people doing it and other people doing that those patterns and I don't know for me that's how I learned. Just seeing other stuff do it and do it myself. So I think we have four more minutes. So if you guys have questions, I can answer questions gladly. I can answer questions even afterwards, I don't mind.
Speaker 1: Yes.
Speaker 2: I say that the
Speaker 1: Yeah, so you can trick it, yeah. I think yes. So yes. So There are some things that I want to showcase that basically the LM is capable to do many stuff because at some point I think I've explained this project to someone Uh and he said like uh how many if L statements are there? I was like, not that many. I mean the LN does most of the stuff. But if you want to restrict some of the stuff, I'm pretty sure you can, yes But then uh you would have to implement some more code for it. But yes, I'm pretty sure it's possible. Um but also I would say with some fine-tuning where you also include examples where the user says that he's lying and
Speaker 1: stuff like that you can make sure that uh it won't hallucinate like this. So yeah, I think it's a it's a big deal that this is a base model and it's quite reliable for this task. Any questions? If you request the page, now what I if you look at the views at pi at the index you will see that I start fresh on page load, so I delete everything And that's why deletes because otherwise you would see everything else. Like i if you didn't actually let me see if I don't do all this Actually no I have to do some other stuff. But if you not delete everything it will
Speaker 1: remain there. Sorry. Yes, that yes, that is true. So if you wanted to have previous messages load, all you would need to do is when you go here Is say checked for messaging, check messages And if and four and basically just add chat message with the role of message at roll and whatever But uh and yeah, all we need to do uh is that and you would have to also pass the chat messages that you get from the session, like not reset them, and then it would work
Speaker 1: Connection errors. Mm-hmm. Mm-hmm. Oh, the thing is HTMX was actually built with a very uh nice approach if you want You can use HX Boost. Some of you might know HX Boost. And then so for example, no For interactive chat chat like this, if the internet goes down, it it won't work of course like you can't uh or
Speaker 1: I'm not sure if I understand correctly. Uh the thing is sorry, go ahead. Oh, okay, yes, yes, yes, yes, that is very possible. For example Uh I c I think I've edited an example of this. I hope it's if I hope I've understood you correctly. Um if I go to the or milestone milestone. I've added the case that if the user did not add it's uh his uh API key It's not here, sorry. It's in the second branch.
Speaker 1: So for example here, I did uh uh try to create a completion, but if it gives any kind of error, for example block authentication error, that was would the that would have been the case if the user did not add his API key. I would just add set the growth API key environment variable and reveal the Docker image to the chat messages and then render the chat message with that message in mind. So you you can handle error cases, but you have to Uh is is this what yes? Mm-hmm. I I'm not an expert at HM HTMX either. I've used for
Speaker 1: my projects, but uh there's still lots to learn. What I know is it has a very extensive list of events that you can listen to and see what's going on through the process like before the request, after the request, after the swap so you can hook into many things and see what's going on and all you need to do is I think that's like an extension it's called hx debug Let me see tricks debug and then if you add it to a specific element you would see in the console everything that's going on during that request so yeah Anything? Uh anybody else? I
Speaker 2: don't consider tools what you're doing with
Speaker 1: uh which to
Speaker 2: a function call
Speaker 1: Yes. Actually it's funny. I would have I like that would have been the most obvious way to choose like to you mean to use instead of JSON mode? Yes. So I did It should, but I don't know why in my experience with this exact project using the tools instead of JSON mode Did not work as well. I can't explain why. I don't know. But I did try the tool use, yes. And i i i instructor does enable this. Like if you if you go to the client Um actually it's in the thinking simulator branch. Second. Uh here. Oh here. Um no
Speaker 1: Here. So it has mode instructor mode JSON, which if you look at this, it has many different many different options like function protools and they it has so many because uh it actually enables you to use it with Google Gemini with Cohereal Entropic. So the guy that made this library, he tried to include as many possible LMs into his project. Which is really cool. Like props to him for making this
Speaker 1: I hope I haven't done a terrible job at that. I just so if you look at the URLs. py here, it does just add that URLs. And then If you look at app, this is like the Ninja Part Ninja API object, which you add URL URLs by using this decorator. App which expects a GET request and sends it for the index page. So that's how it that's basically you declare the URLs just next to the views. There's actually an attempt for that. If you look at the Django Ninja 's uh uh documentation, it wasn't yet implemented, but there's there's an attempt to do it.
Speaker 1: But
Speaker 2: is is this uh instructor or something you would generally recommend?
Speaker 1: If you want to get structured inputs, like it's been a very useful tool for me. I think there are also other tools that do the same thing. What I liked about this is that he made sure that um your I so he made it possible so that you can just wrap any client, be it Groc or OpenAI. And your IDE won't uh yell at you because uh suddenly you had an argument that I don't know. He made sure all the internals were alright so that as a user as a developer as you use it uh it's it works well. So It's pretty cool. Yeah. I'm sure there are others as well. The
Speaker 2: question like how we are sending all the messages. Is there a way to kind of make it a bit like smarter that's or the the system keeps the context in the expandable or is it a compact that you have to return on the side?
Speaker 1: This is how every LM works. Like this is this is how ChatGPT was as So far this I mean there are many interesting attempts like there's so much development in the area of LM and people looking for more ways to just make more like Google has like one point one point five million tokens contacts now, which is insane. And there's many things like uh some pretty much reduce the uh previous conversation into some kind of the s compressed state of those tokens well it still keeps them I don't know I'm there's a lot to learn Any other oh and there's actually I think isn't your your library Django Ninja Crut isn't it a class based approach that uh the other guy was looking for?
Speaker 1: No
Speaker 2: Ninja Extras.
Speaker 1: Oh okay. So there there is one? Django Ninja Extra? Okay. So there's a question. There's a answer for the person that has. So there's there's actually a library called Django Ninja Extras. It's her party, right? From a guy and it does include some niceties for Django Ninja, including hot uh objects classes, yeah. Um so I guess that was it
Speaker 2: This was using Starlet.
Speaker 1: Yes, generally runs on Starlet, but
Speaker 2: starting for the L that's on it right now. It just you chose to
Speaker 1: No, it doesn't matter. Yeah. So uh I would love if you would scan this QR if you want and give some feedback about how this went. Because it will make it let me know whether I should tr keep trying or just stop wasting people 's times
The application keeps a list of messages containing each message’s role and content, appends new messages to it, and sends the full history with every new completion request so the model retains the conversation context.
Discussed at 7:46Instead of submitting a normal form and refreshing the page, HTMX sends an asynchronous request and receives an HTML partial from the server. That partial is inserted into the existing page, creating dynamic behavior without JavaScript-heavy frontend code.
Discussed at 10:07The model can return structured fields describing the action, which the application converts into HTMX client events. Those events trigger existing HTMX requests that update the interface, allowing the LLM to control predefined UI actions rather than arbitrary behavior.
Discussed at 37:11Wrap the Groq client with Instructor, select its JSON mode, and pass a Pydantic model as the response model. The result is returned as an instance of that model, so the application can use typed fields without manually parsing and validating JSON.
Discussed at 38:41A short reasoning field can make the model’s decisions more reliable by having it explicitly consider what the user requested before taking an action. Without that restraint, the model may trigger available actions impulsively even when the user only says something like “hi.”
Discussed at 42:45Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025