3D Files with Wagtail
Published July 19, 2024
This video is from Wagtail Space US 2024 in Philadelphia, Pennsylvania, USA.
After eight years and 10,000 pages, the content stored in a Wagtail site can become unwieldy. What blocks are we even using? Where are they being used? Maybe we've identified a pattern in our content that we want to remove, where do we find ones like it?
At CFPB, we've been going through the process, auditing the content stored in our Wagtail site. We'll talk about doing that audit, tools we've created to help, lessons we've learned from it, and how it may have changed our approach to both the how we think about our content in terms of both schema and admin usability.
💻 Wagtail is the easiest open-source Python CMS to use:
Install the demo and start building your first site in 10 minutes: https://wagtail.org/get-started
📹 Related Videos To Watch Next:
â–¶ Quick Video Tour of Wagtail CMS 6.0 https://www.youtube.com/watch?v=_Vg_lPMipcQ
â–¶ The Latest on Wagtail AI https://www.youtube.com/watch?v=4zfs1u4Vy5Y
▶ What’s New in Wagtail CMS 6.0 https://www.youtube.com/watch?v=2AxLFyOFjQo
Wagtail future proofs your CMS system, as it’s open source, continuously updated and built on Python, one of the most popular global programming languages, used widely in machine learning and big data. So you’re always ahead of the curve when it comes to CMS platforms.
Wagtail is the #1 choice for accessibility, is scalable and most importantly, secure.
👉 Get started with a FREE Wagtail CMS TRIAL: https://wagtail.org/get-started
and see how easy it is to build a website that works for you.
📊 Read why Google, NASA, and the British NHS, are powering their digital estates with Wagtail: https://wagtail.org/about-wagtail/
🎥 More Wagtail Videos: https://www.youtube.com/watch?v=cne2kxemMAQ&list=PLfwZ-fob20cPvSQ_v1hkjto8BAPN21tLJ
📣 Follow us on social:
#WagtailCMS #Django #WagtailSpace
The CFPB’s Wagtail site grew over eight years into a complex system with more than 30 page types, deeply nested StreamField blocks, legacy WordPress content, duplicated concepts, and a gap between content owners and the people publishing their work. To understand and reduce that complexity, the speakers built audits that report page-type and block usage, recursively inspect nested StreamFields, locate raw HTML hidden in text and rich-text fields, and export results for analysis. These audits revealed unused options and recurring workarounds, which led to simpler editor interfaces, proper support for anchors, icons, and expandable content, and data migrations that removed unsafe HTML-unescaping behavior. Their main advice is to trust past decisions in their original context, use accumulated content as evidence of real user needs, and address a large legacy system incrementally with repeatable tools rather than trying to fix everything at once.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: So without further ado, I'm going to turn it over now to um to my friend Will Barton and Chuck Zebey and Lander. We're going to talk about auditing Wagtail content and how to do it. So please help me welcome uh Chuck and Will.
Speaker 2: All right, thank you all very much. Um First, uh we have to get this out of the way. So we're from the CFPB. Um we're a US government agency. We're set up after the 2008 financial crisis to implement and enforce federal consumer finance law. and ensure markets are fair to everyone. Because of that, we have to give this disclaimer that we are representatives of CFPB, but this does not constitute any legal interpretation, guidance, or advice. Any opinions are ours alone. So the other thing as CFPB employees, you may or may not be familiar with some of our packages, but All of the code for our Django and Wagtail website and libraries that we've created around it are open source and they're on GitHub.
Speaker 2: And we work entirely in the open there, which I think is a good thing for a government agency to do. But you know, it's for better or for worse. You can see some of our successes and failures as developers. So uh I want to start with a little bit of history. So we have been a user of Wagtail for a little over eight years now. Um We adopted Wagtail around version 1. 1 in September of 2015. And the site that was based that based on Wagtail was launched to the public in May of 2016. It had to accommodate data from a prior WordPress site that was really hastily assembled when CFTP was established in
Speaker 2: 2012. and grew unwieldy with huge numbers of custom fields. Wagtail let us streamline that kind of page architecture for content managers where we could have a limited number of page types with potentially many fields that would meet various different needs all in string fields and content managers would have options while also like limiting the number of unused fields they had to wade through in the admin. You're all presumably familiar with this use case to some degree, and like that was the theory. With this powerful new CMS and all of its flexibility, we felt like we could do anything. You know, as designers, developers, content managers, uh
Speaker 2: when the Bureau identified needs that the public had that they wanted the website to meet, uh we could make them in as many ways as we as we could. You know, if we need a new block or five added to the stream field, it's done. We need a new page type for a particular type of content. We need a new stream block that's nearly like another one, but not quite. For use on a single page type, that was done. And so you can all see where this is headed, probably especially if you were uh here for uh Michael's talk yesterday morning. Um so the question, first question is kind of like, should we have used the power to do all of this? And I think the answer is yes, because I don't want
Speaker 2: the talk to evolve to devolve into like condescension of our past selves, and I'm not here to tell you any of this was mistaken. the time um and to learn from it and limit what you implement. I I think the first step is to trust the decisions that we made at the time that they were the right decisions to make at that time. You know Future U might make different decisions based on extra context that now you doesn't have, and that's fine. But that doesn't mean that we didn't end up with a really complex pile of code and content.
Speaker 3: Which I'm going to walk you through a little bit right now. So just for a little bit of extra background on my my perspective. So my my role in our team is as the web content manager. Which I'll explain a little more what that means in a second. But basically, where I see this complexity is in the size of what currently is in Wagtail. So this is a snapshot I took a couple weeks ago of what you currently see when you log into our dashboard and the CMS. This of course doesn't include Um some of the pages we have that are generated dynamically, they don't count as individual pages for Wagtail. We still have some static HTML pages that are accessible and look like They would be in Wagtail, but are not because they're legacy content or things that are generated as pure pages from snippets or things like that. But there's still plenty.
Speaker 3: That includes a lot of page types, as you might expect. We have over 30 of them and the use cases for them range from really standard use cases like blogs. We had our newsroom page type, which is used for things like press releases, director statements, things like that. We also have minor derivations of page types. If I tried to give you the nomenclature, I'd have to put on the whole rabbit hole of our naming standards for some of these page types. And we have the a smattering of one-offs that are used to build a particular type of page that was needed somewhere once. There's also legacy content, as well mentioned, that predates Wagdale, the content we pulled in from WordPress, and any content that had existed on the site. um pre -2016. We will show you a little bit of what that looks like in practice in a bit.
Speaker 3: The thing about it is like being this old or being this sort of um built upon the legacy of the content from its very inception is that you get a lot of accumulated institutional knowledge and difficulty maintaining consistency with that. So there's a lot of arcane nomenclature or standards around the naming of things like rulemaking or or enforcement actions or compliance. That The person who was doing content management on the site in 2017 might know, but then when I first came in in 2020, I had to learn it all from scratch. Because really the only place where it's documented is the website itself. The complexity here is driven in part also by a weird, this is probably not entirely unfamiliar to some of you, a separation between content management and content ownership. By which I mean the people who write content that goes on our website are writing it in Word documents and sending it to my team, which is pretty much just me and one other person.
Speaker 3: to publish on the web. And so there is a whole layer of translation that is happening there, or a whole layer of loss in translation that is happening there. I mentioned the newsroom page type That has the way we have that built in Wagtail, it is a single page type of newsroom that has categories that can be assigned, press release, statement, etc. That is not how the writers of the content think of it. They tell us to say a press release, publish a press release. They don't think of those things as being the same thing. And while that's usually not a problem, it can lead to some really weird discrepancies in how you are understanding the content as an owner and seeing it at Wagtail and as they understand it looking just at the website. And that sort of problem only accumulates further. over time. What you also get over time is weird almost redundancies. So these are two different pages with two different definitions of the same term.
Speaker 3: They are probably functionally identical in many respects, but the problem is that they're serving different use cases. One of these is coming from a glossary terms page, which is meant as a reference point when you are in other types of content. And one of them is specifically on the section of our site called Ask CFBB where a user would specifically have the question, what is a reverse mortgage? Which implies they are a user who is interested in one, needs more information about one, and so we are speaking to them in a different way. So these are not identical enough to be functional as a snippet, for example. They need to be different for different use cases, but at the same time, if something changes about the definition of one. it's almost certainly going to change about the other. So there is an asymmetric, I need to keep content synchronized when it also can't be synchronized. That again only grows the larger your website becomes. Just give you a random example of the sort of problem we have to deal with because of all this.
Speaker 3: Back in 2022, we were given a request to change some language on the website involving the Department of Housing and Urban Development, HUD. So HUD Among the many things they do is that they approve housing counselor housing counselor agencies, which are agencies that can provide you a counselor who can give you advice um on whatever housing issues you are running into. We had some language on the site that referred to HUD approved housing counselors, but the thing is that that's not actually legally true. HUD approves the agencies, they do not approve the individual counselors. So we were asked to go across the website and change all the language that refers to HUD-approved counselors and change it to HUD-approved agencies. And the thing about that is That's changing it everywhere on the site regardless of page type, regardless of
Speaker 3: whether something is stored in a snippet, stored as part of a blog post, stored as part of some weird legacy content that everybody forgot about because it was last updated in 2017 but refers to it repeatedly. That might be stored in a plain text field somebody Over here, a rich text field, over here, a snippet over here, in a call-out field or something like that. But the people who are asking us to do this don't care about any of that. They're thinking of it as the website. the language needs to change. So we have to take requests like this and distill them with a website of this scale into something coherent as something you can do on a website like this. Which obviously leads to the question, where do you even start?
Speaker 2: Yeah. So this is a daunting problem, right? So you've got all of this complexity built up over years and years. And we don't want to solve specific content issues. We want to understand what the nature of that complexity is so that then those content issues can be more easily resolved, right? So like how do you even start interrogating all of this to know what patterns might exist? So we identified a few specific kind of problem areas that we wanted to get information from. And that's the most of the rest of this talk is going to be about how we extracted information to try to then make decisions about. how we we go about this. So we have a lot of page types.
Speaker 2: You know there are questions we could ask about that. So how many do we have? How many of those are only ever children of specific pages, that sort of thing. We have a lot of blocks available on all of those page types. How many? How many are available but unused on some on certain page types? We have stream fields that have struck blocks that pull in blocks from all over the place, right? And we have some places where raw HTML was entered into content fields because the available blocks didn't quite do what someone needed them to do. And yeah, we'll come back to that one. So let's start with page types. This one's somewhat simple because Wagtail now has the built-in page type usage report. It's awesome. It tells us exactly how many pages we have of each type.
Speaker 2: And we can see this is the sorted by maximum, obviously, from the top. We can also see we have these two legacy newsroom and legacy blog page types. Which Chuck mentioned. We'll come back to those two later. The other thing this page, this report shows us, if you scroll all the way down, is The pages that are only ever used once. And something that you can notice about some of these one-time use page types is they're almost all landing pages of some description. Which suggests that you know maybe our like for example our newsroom page type uh could be limited to being a subpage of the newsroom landing uh page.
Speaker 2: Um it also suggests that like Probably it's fine that these are one-offs if they are landing pages that then uh like have a really specific use case and have child pages of of other types. So yeah, generally the the pages themselves or page types themselves are probably the easiest thing to audit. We have this report built into Wagtail. They are also Django models, which means we can easily construct our regular like Django query sets around them to look into them without a whole lot of effort. So that gives us some basic information about the pages that then we can take and we can act on, possibly in combination with others. So what's next?
Speaker 2: Blocks. We have a lot of blocks. When we start considering which blocks are used on all of those page types and in which stream fields, we get into the weeds of stream fields a little bit. Like I said earlier, like stream fields are one of the major draws of Wagtail to us. You know, they they permit a lot of flexibility on a page or four-page content without it being a UX nightmare for anyone in the admin. So I want to look real quick at an example of how we're using them. So this is a part part of a page that we have to help renters who may be struggling to make payments. This is In the admin, this is part of our info unit group, which contains info units, which that was
Speaker 2: , which can have images, body text, links, and other optional fields. The Wagtail admin is great about uh displaying uh hierarchical sort of oops. Um Hierarchical relationships between these that the image upload for an example and the body clearly belong to the info unit, which belongs to the info unit group. And again, in in recent versions of Wikil, obviously the divider lines like they they highlight as you go up as you mouse over the different levels, right? So what does that look like in code? We have our info unit group, which is a struct block It has chat blocks. The info units
Speaker 2: are a list block that contains info units. So the info unit block itself is also a struct block. that uh includes our image, our our basic image block, rich text block, and hyperlinks. So That's our hierarchy in the Wagtail admin, our hierarchy in code. What does it look like in the database? As we'll probably know, screen fields. are JSON columns in the database. And when we look at the database content field, this is somewhat somewhat edited for brevity with it on the slide. But we see something like this, right? It's a JSON data structure. That 's relatively expressive.
Speaker 2: So we know that that's how stream builds are stored. So if we ask how many pages are using our image basic block, which again is is the if I go back here, that's our like basic image chooser block that we that that we use most everywhere we need images. How would we go about answering that? Like our that This block is not only in our info unit struck block, it's also in our featured content block, in our content image block, many, many others. Some of those are also included in struct blocks that include all the other blocks, right? So what if we want to count all of them?
Speaker 2: So both Django itself and Postgres VL have some really nice JSON support built in. For Django, if we know the JSON path, We can query into it directly, which is awesome, and I love this. It's a relatively recent feature, and I don't remember which version. The problem is though that with stream fields, we don't necessarily know what the concrete JSON path that we want is for that image block, right? It can occur at various different levels in the JSON because The stream field schema is not concrete in that way. So what we actually want to know is what the path is and then how many times it's used. And all and So yeah, Postgres
Speaker 2: also fantastic JSON support built in. We could build a temporary table. Plenty of lateral joins, all matching blocks within the JSON, query that. And then maybe we could wrap that in a Django query set. And maybe that gets us somewhere we wanted to go. Unfortunately, while that code mostly worked in Postgres, we could not get that to work as a query set. And we got stuck on this notion that we had to do this in the database, right? I know this is this is a feeling that we all probably have. I hope I'm not unique in that. If I am, then this is a me problem. But um But it should be in the database because like that's the way it's performant, right?
Speaker 2: We stepped back and started thinking about this. Okay, we want to do this once right now. Maybe again, maybe we want it to be repeatable, but it doesn't have to be like a few minutes of human timescale is okay, right? Um and We're a team of Python developers. It makes sense if we try to do it in Python because someone after me who knows Python can understand it and they don't have to have like intricate knowledge of Postgres, lateral joins, and JSON V support. So fundamentally there are two things we want to know, right? We want to know which blocks are available on each page type and which of those blocks are in use. So we ended up with two recursive functions to walk first to the stream block, gather all possible uh paths
Speaker 2: to uh available blocks themselves and then to walk through the stream values on the pages to find which of those are being used. We used a data class to hold the results. Don't worry about This, I know it's small on the screen. All this is open source, and I'll give a link to where this exists. We wrap this in a library with a management command that we could run to run this as an audit to get our block usage count. And so yeah, that's available at this URL here. Um, and I'll dump the URLs, etc. in the Slack. But we published a library that includes that management command, lets us just get a CSV of block usage. And we could, we added
Speaker 2: The ability to filter based on page type. So if we just want to look at blocks that are in a specific page type's stream field, we can do that. We also decided to use a library called Query-ish, which is really awesome for creating Django query set -like objects that let us express this like it's a query set with the idea that maybe one day we can wrap this in a ragtail report. That is not does not exist yet But um so that's also part of this. So the result of this we output to CSV, we can do some filtering in Excel, um
Speaker 2: and eventually get to seeing where our image basic block is used. There are a couple of things to highlight here. One, it's available on six different pages, six different places on our learn page. model, but it's not used at all in two of those. So that's interesting. That gives us something that we can actually dig into and start looking at and see, is that okay? Can we remove that as an option and simplify it for our content managers? The other thing that I want to note with this report is the path column is um As close as we can get it, we haven't found any examples where it doesn't. It matches the path, the the path that JSON
Speaker 2: that um Wagtails new Streamfield Migrations Expect for doing migrations on JSON paths within the stream field. So hopefully that's useful should we need to do any migrations based on this kind of data. Hopefully it makes it a little bit more straightforward, I hope. So that's good. That was exciting to us. I can tell the room is thrilled looking at small screenshots of Excel files. But we were certainly excited. We could pull useful information out of our stream fields and our deeply nested and somewhat incestuous blocks. And that gives us things that we can potentially act on. We can do that audit, we can figure out what's going on.
Speaker 2: And that's cool. That's what it's It's all about.
Speaker 3: Yeah, and something I want to emphasize about this that we'll talk about more in a moment is what's happening here is if you've ever had the experience of um Being a developer being asked to add a block to a page type where you don't think that block is going to be used, but someone's like, well, maybe we're going to need it there. You're like, all right, fine. And you suspect it's never going to get used, a really good way to test that is just to have your website run for about eight years and then chat. Because what Will just showed is essentially It's sort of a it's it's in some ways the benefit of having a website running this long is you kind of have a built-in user testing function here where you get to see in practical case, well we have a lot of examples of that page and we don't use that block. We don't need to have it here So this audit was really valuable for us with a website with this much legacy where it can be really hard to sort of um
Speaker 3: manually take a look at every possible variation and say what makes sense and be very clear-headed about not just the zeros, but also the ones and the twos, be like. Okay, you used it there, but did you really need to use it there? Can we maybe switch this and standardize it? So this is like it's a really nice Where a lot of this talk is built on the premise that this is a bad thing, that we have this ancient website that's creeping along, there's some really fundamental advantages to having this much data to use about the usage of your website in Wivetale.
Speaker 2: Alright, so we've got our page text. We have some understanding of what's there. We have some understanding of like how our blocks are being used. What else was there? Right, there was that one. So sometimes the CMS doesn't do what our users need it to do, right? Um maybe we don't have a block for that, maybe the block doesn't have the right options. Maybe we don't have the capacity to add that option right now. It might be way down the priority list. So what's a user to do, right? So we were able to identify at a glance a few places where raw HTML was being used. We have places where we include Wagtail's RawHTML block.
Speaker 2: intentionally, uh we know we might need to use where we know we might need to use raw HTML. As a little side note, this has also become a way for us to embed re like one-off React components, which is Not the best approach, but that's like a whole other talk. But it's reduced a lot of friction to to introducing that. We had some raw HTML in text fields, which is not great. Yes, that's some Roll HTML in our rich text fields. Oh boy. Yeah. So what are some examples? After all, we clearly have some users who have needs of the CMS that it is not meeting, right? So what are they?
Speaker 3: Yeah, so to start with, I'm gonna go back to explaining what the legacy newsroom and blog post uh content was. So when we migrated um out of the WordPress site that we had had pre-WagTown migration, We needed to import all of that content, all the blog posts, press releases, et cetera, into the new site. And the solution for that was to drop it all completely as raw rendered HTML into single content blocks on this page type. So this is what basically every legacy blog or legacy newsroom page looks like on the back end. On the front end it looks like a blog post for everybody else. And on the back end it is as far removed from what you expect when you open an editor and a content management system as I think is possible. You're going to notice fun things that the further you go through this, like classes being referred to that don't exist anymore in our design system specs.
Speaker 3: uh like specifications and uh made enforced in html that need should be getting rendered on the front end um all the sorts of stuff you would expect when you've dropped a giant mass of html in but this was you know the necessity of the time that has been a then legacy that we have had to carry through since then. We'd ask to try and remove some of this. Work in progress. Um, so in the meantime, we just have to figure out how to mitigate some of the problems that arise with this.
Speaker 2: In our plain text fields, we also see examples like this, where we've got, you know, we're overriding a heading level and adding some padding. Maybe we're adding strong and emphasis tags. And um yeah do we we have some needs that we're not meeting, right? I also said we have raw HTML and rich text fields, right? What's that look like? So we have users adding anchors as an example to rich text. And so for those of you who might have a more familiarity with DraftTale and how rich text works, you might be wondering how how does that actually
Speaker 2: work. Because it doesn't get stored. It will get converted to HTML entities, right? So the greater than will become ampersand GT, etc. And yeah, in the database, it looks like that. But that doesn't answer the question. How does that work? So Way back in 2015, when we adopted Wagtail, we made a faithful decision, fateful decision when rendering stream child blocks. I suspect that this was before a lot of the deeper rendering options for Streamfields existed. It's definitely before the include block template tag.
Speaker 2: So this code existed in our render cycle until after we did this audit that the presentation is based on earlier this year. This is how HTML entities in our rich text field became actual HTML. That's how our CMS users could put any HTML they wanted into our rich text blocks and it would render. So yeah. At that point like so maybe the problem isn't just span tags that are anchors. You know, we we can provide anchors in the rich text editor, that's fine. But what else is even in there, right? How do we find out what's in there? And in which rich text blocks, which
Speaker 2: they can be in it many different places on a page. So we'd already solved that problem of walking through blocks, identifying specific paths, etc. Maybe what we need to do is use that pattern again, make it a queryish query set like that we we did, like we did before, and maybe add some searchability inside those fields along the way. We probably want to start by filtering our page query set using Django's iRegex filter. On that field, we can search for particular patterns somewhere in the field that we're searching, right? And if the field is a stream field, then we could walk through it like we did with our block usage audit.
Speaker 2: in a recursive function and maybe we use a data class to hold the results. And again Maybe we want a regular expression to try to match HTML hiding in entities. We'll gloss over that a little bit. Hopefully we we all know how Dangerous, uh dangerous, but how frustrating regular expression matching with HTML can be. So yeah. We have our page query set. We can do that. We can filter it based on page model and field. We can give it a regular expression. And we can run it as a management command. And this also exists in that library that I was mentioning earlier.
Speaker 2: So Through a lot of iteration , we got some results where we're matching with I think a few false positives that our regular expression matched, not too many, which was good. Again, one per row of a match rather than a page. So um For every time that we matched our regular expression looking for kind of like quote-unquote raw HTML, the the unescaped or sorry, the HTML hiding as HTML entities. We could get the exact location. In this case, we have like our um oh that works there too.
Speaker 2: Awesome. So uh we have the stream field path. We also have a result path which gives us an index within stream field. So some of these are list blocks, right? So we want to know that it's in the fourth um member of the expandable groups fody, right? Or in this case the sixth Expandable group, expandable, the first item, content, and the zeroth paragraph. But that way we can open the page up in the admin and see what in the world is this thing that we've got, right? Which is useful again for finding those patterns so that then we can develop a strategy for remediating them.
Speaker 3: Yeah. So there were a fair amount of one-off use cases that were easy enough to manually remediate. But uh broadly speaking, there were two broad use cases that we handled as a result of this that were like immediately easy to see in the data, which was that anchor uh link use case that uh Will mentioned, and then the use of inline SVG icons. Um We um used we implemented Wagtail Drafttail anchors to add that anchor support properly into the rich text editor, and then we migrated this pattern of text to use it. So this is now what that looks like which is so much better from an actual usability on the back end situation for content managers to understand what they're doing
Speaker 3: Basically the only problem left that we haven't fully solved yet is the SVGs, but we are actively working on that. We're getting there. But everything else, every other use case of HTML that was being used in this fashion in Richtext fields has been remediated. And something I want to I want to come back to something I said before about this. So when I mentioned when you've when we found the cases of like a block is never used in this particular combination, that's essentially an example of sort of like user testing by accretion over years. This is another example of that. Like Will said, what this is demonstrating is not people are trying to break the CMS. It's people trying to do stuff that they need to do, trying to create patterns that they either think should exist or are supposed to exist per some other place and they're forcing it into place. So sometimes you might look at something and say
Speaker 3: That should never a heading should never have a strong tag attached to it. So you just remove it. And this was actually recently we updated our own design guidance so that there's more bold emphasis um uh the font weight was changed for headings so this became sort of obsolete to begin with so those cases sure you get rid of them. Um likewise when we're sort of brute forcing headings into label fields Why not just put a heading field there that solves the problem by actually creating the pattern as in the way that the person clearly thinks it was meant to be uh built? Um We had cases with the SVG icons of SVGs being placed in our label fields because somebody wanted a heading with an icon next to it, like maybe a question mark next to something about getting help. We solved that by adding an icon field next to the label field in these cases
Speaker 3: because again, that was that was used consistently enough. and on fairly significant pages on the site, that clearly it was a pattern that people wanted to be able to have access to. And it's something we could easily add programmatically so that it's being added consistently as well. We had some cases of people forcing boldface on certain links on these link fields. And when we explored this, we found that the use case actually made a lot of sense. This would be the it's usually The last link in the list of links where it was meant to be the broad like see the rest of this information. And it makes a lot of sense visually, and this is in consultation with UX people and our designers, as well as the people who wanted this page built in the first place, our content stakeholders, that that made sense to the users. So we added is link boldface as an option. You know, it's it doesn't have to be it wasn't so much about making this the perfect UI experience as it was about accommodating the actual needs that we were seeing in practice.
Speaker 3: Um and again going back to Not only adding the icon label, but adding all of the other things. This was an example of people trying to not just add the SVG, but if I could scroll through that endlessly long single line of text, you would also see some additional padding and sizing things because they wanted these to be bigger, expandables, which was something I think is in our design pattern library, but was not actually doable. in Wagtail because the expandable didn't have that functionality, so we added it. So this is again an example of how The the reviewing the data and looking at the data allowed us to basically meet the needs of our users, of our content users, of our content owners, and create a back-end in my tell that is actually serving what they need to do so that there don't have to be solutions like
Speaker 3: Unescaping HTML.
Speaker 2: Yep. And so several remediations and data migrations later, we removed that unescape that was hiding. And then quickly fixed a few minor things that we missed. So like I said earlier, that was a daunting problem to consider. Just that problem of how do we even know what patterns we have that we want to start. addressing. And there's a quote I think that's usually attributed to Desmond Tutu. The problem with that is that I'm a vegetarian. Which, you know, in a way that does still work for this talk.
Speaker 2: It's like how do you approach that massive pile of content? You just don't. But that's not what we're here to say. So You know, do it a little bit at a time in small pieces and hopefully thinking about your users along the way and creating some useful tools for ourselves maybe and for others. If we need to do it again in the future, which
Speaker 3: certainly we probably will.
Speaker 2: Yeah. So yeah, thank you. Uh does anyone have any questions? I think we're we have five yes.
Speaker 4: future based migration. Um I'm just curious if you ever ended up using that for the future based migration technology.
Speaker 2: So the question was if we ever used the stream field path in data migrations uh in our our block usage report? I honestly don't remember. We have done a few data migrations around this that we did use a path. I'm not certain it was a path that we took out of that report though.
Speaker 3: I think we discussed it and the possibility at the very least.
Speaker 2: It's a good question. I I do hope that it is that it's useful, uh, because sometimes discovering that path Discovering what that path is for providing previous stream field data migrations that I 've I've had to do is not Always intuitive because again, some of our blocks are very deeply nested. And sometimes I think we've eliminated all cases of this, but that there have been cases in the past where we have a block, a struck block that has other struck blocks, and eventually you get to another struck block that is that original struck block. So theoretically you could have infinite sort of uh struck blocks all the uh of that original block all the way down. Um
Speaker 2: but yeah that was the motivator for including that. I don't remember if we actually used that. Yes.
Speaker 5: I mean how does this interact with Lartel and Twitter?
Speaker 2: That is a very good question. Um yeah, so uh Kind of two questions. First was uh how this interacts with Wagtail inventory. Um and the other was about um stream field migrations and the patterns that we were using before Wagtail introduced their own. Um so with Wagtail inventory It doesn't. This is all kind of independent of that. It's Vagtail Inventory is very useful for identifying Like I have a question about this particular block, where is it used on which pages? And then going in and looking at at individual examples of those, right?
Speaker 2: And I I think it's it's still useful for doing that. This is it it doesn't give us broad trends across, again, our our nearly 10,000 pages. And that's where we wanted to do something a little bit more useful. Yeah.
Speaker 3: Yeah, I just I think it to me it's it's essentially that the Wagtail inventory uh versus it just sort of solving the problem in an inverse relationship where lacktop inventory, you know what you're looking for, you search for that and you see how many examples of it are. One of the big questions that we had, and the reason why this came was like We don't know what all the combinations are. Like especially 'cause it can get so nested and so convoluted that like we need what we wanted was I just want all the possible variations in a big data set that I can then come through rather than coming in already knowing what I want to search for. So I think like since the path in is sort of fundamentally pretty different, I it it made sense to sort of tackle this as a separate problem. Yeah.
Speaker 2: That's all like none of this is to suggest that Wagtail inventory is going away, right? It's still very useful. The stream field migration question. Yeah, we we we've been again. With our really complex stream fields, we've had to tackle migrations long before Wagtail had any built-in support for doing stream field migrations. We have migrated away from those patterns because Wagtail's built-in tools are better at this point. Yes.
Speaker 6: Yeah, so um I also have a site that we've been running for eight years and stuff like that. and are about to do a content audit of it. So this was highly relevant. sort of the FAQ version and the the um you know the I guess the policy or the just the actual statement. That's that's really a case I kind of live in fear of. And I wonder if you found any kind of shortcuts for identifying those kinds. of out of sync um attempt or just the old fashioned way.
Speaker 3: So we that's a sort of a long it's like a long standing I feel like just general problem of um Especially when you have a site of that scale, I think the expectation content people tend to have is like, can't you just do like a find and replace like Word has and just push a button? What what this allowed us, what some of this work allowed us to do is sort of narrow down the use cases we had to check. And that came in tandem with um, I think it was actually in the in for the case I showed. It was actually the content owners for the sections of the site where that would be most affected had run some external crawler tools to give us at least a starting point. And from there we could then check patterns for usage. So If we're seeing from their results, because they they won't know. They might say it's in the sidebar all the time. But of course this the sidebar is a thing that exists on the front-end pattern that is not represented necessarily
Speaker 3: because something is in lagtail. But it gives us if if we have these URLs, we can say, okay In all of these, it's a manual content block that's being added to the sidebar. So let's just check that, like let's let's narrow our focus to that particular like pattern or variation and make sure we've caught all of those. So it's sort of a tandem thing, but I think there is a level to this. And we could do probably a whole other talk about the crawl. We've built an internal crawling tool that essentially crawls based on the rendered HTML to work in tandem with LagTail essentially to solve that problem of the translation gap between what a user who isn't in MikeTel thinks the website is and what it actually looks like for users of MikeTel. That tool is essentially to for those people to say, I want to connect everywhere. I don't care if it's in a rich text field. I don't care if it's in a heading. And then we can take that data and combine it with our ability to do sort of light scale auditing
Speaker 3: of where that lives in the system and sort of nice and match. So it's not beautiful, but it is definitely better than more tooling like that you have that people understand and and and know what they're coming with.
Speaker 2: Uh yes.
Speaker 5: Thank you for giving the talk that I was attempting to give.
Speaker 2: Oh, and yours was yeah.
Speaker 5: So I've hit this problem a lot. I just went through a project that had 100,000 objects, like 20,000 images so we just when it was going WordPress we can do the whole thing. Um one of the problems that I've had with supplies uh is they don't see the problems that you're noting here. They don't know why these are issues. And it's a really tough sell. Um did this come Did did you folks decide to do this because it was just a smell and you had a feeling or were you hitting the problem so they could point to you and say this is
Speaker 3: a So that's a really good question. And I think it like it's it's
Speaker 2: uh
Speaker 3: Oh yeah. So so the question was um Go ahead if you have a way.
Speaker 2: Yeah. Well um sorry, I had it.
Speaker 3: Sorry.
Speaker 2: Yeah, no. The the the question was uh like this is the sort of thing that we notice on the back end that isn't necessarily obvious to people that are working on the content work that like look at the website and it's like why is this a problem and how do how does that work get prioritized? Is that a fair something
Speaker 5: yeah there's We see the problem, but to them it's fine. So it's fine then.
Speaker 3: What happens when it's a it's a a back-end problem that to a front end user they would never notice that it was a problem, which in a lot of those cases is exactly what it is because rendered HTML is rendered HTML, they don't care.
Speaker 5: I'm also kind of yeah, and I'm kind of curious as like this was a tremendous amount of work having just done it. Yeah. Yeah, it takes time, money, effort, move away from other things.
Speaker 3: Right. So so so yeah, the question is basically like where is the motivation from outside of our own team and therefore the prioritization, budget, whatever you need. So there's a couple different answers to that. One is that um if the right HTML is a is a good example of where you can we're a government agency that has pretty high security standards. It was pretty easy to basically say like Unescaped raw HTML that can be is arbitrary in our backend is an extremely huge security risk because it takes one person, one bad actor, can do some really nasty damage with that. That in and of itself probably would have been sufficient justification. Something we've also pointed to what I have tried to do when having those kind of conversations is Usually when those things are being done by people in the system in that way, it creates two problems. One is that it takes a lot longer to do work that people feel should not take that long.
Speaker 3: So the span tags for anchors. We have research reports that get published on the Bureau's website all the time. We would have reports come in that are 100 plus footnotes. You're asking a content manager to manually type in HTML and to set up all these anchors, and the report is taking A few hours maybe longer than they expected it to. And so then you have an after action report and people are like, why did that take so long to do? And then you can think, well, there's some problems on the state on the content finalization, and and we probably could have gotten the content sooner, but also It's a huge pain to do this on the back end. We had the advantage of recently at the bureau, relatively recently, along with the new the director Chopra who came in. uh with the new administration. We also instituted a chief technologist office and that office and the people in it have been very good at advocating for
Speaker 3: thinking of technical problems as actual problems that need to be solved. So that helps with these two. But to me it's it's like highlighting um Failures in work streams where things take a lot longer than they should, inconsistencies in the end result. When somebody's doing a lot of manual work like this, they're not always going to do it the same way every time. And it's an e it's the easiest way to violate design patterns. We also have accessibility requirements. Once you're violating success once you're violating design patterns like this, if you're brute forcing a spantag and or brute forcing links in certain places they shouldn't be, you're probably not carrying with it all of the um associated 508 sexually 508 compliant, very carefully browser tested pattern library stuff that we pull in as a matter of our fundament library. So it's like, it's, it's always, to your point though, it is always kind of a navigation when it comes to somebody who can't see any of that of trying to explain it to them in the
Speaker 3: It actually is going to cost you more money, it's actually going to take longer, and it's actually going to be a bigger risk than you realize. And that works better or worse in certain cases than others.
Speaker 5: I'm always surprised at how much pain you for willing to go just to get your job. Well and you know that experience.
Speaker 3: For my part as a again, I came in as with very little um developer interface with this website. I came in as a content manager, and my experience in in past jobs doing that work was always like I can get the system to do what I needed to do. I don't really care. And it does really take some amount of that content manager team being somehow integrated or being given visibility into the development team. for those conversations to happen because it's really a you don't know what you don't know. You don't realize as the content manager, oh, this could actually work so much better for me if I just was able to ask. And the developers will be more than happy to do it because they also hate what you're doing. You just haven't had the convert the point where the conversation is happening. So the fact that the Bureau, the team I'm sitting on is a cross-disciplinary team. that has someone who's content manager, someone who's a UX specialist, someone who's a developer. And those combinations helps those conversations happen organically in a way that we can be tremendously useful
Speaker 3: for both for identifying problems and then for coming up with the right rationalizations and justifications when you are you need to sell it to somebody higher up the chain to that the work needs to be done. Yeah I think going along with that one of the questions I had is like And
Speaker 6: you didn't want to rehash past the things too much. One of the kind of interesting things that we saw is these kind of like desired paths were created for like components that didn't exist. the back end that you know after several years maybe it's like obvious well here's how we should write that block and it's kind of an interesting catch point too because like I'm sure most of us are often caught at you know the early implementation side trying to like guess like I noticed like the icon
Speaker 4: is like oh okay it's just the name of the SVG like that works for someone because you had eight years for that
Speaker 3: right
Speaker 4: writing HTML but like you don't always know you know uh what the user needs are there so I wonder if like moving forward like how do you how do you plan to like integrate that into your team and also the candle like I'm sure you still have special cases that you now are not handling
Speaker 3: with
Speaker 4: RTML
Speaker 3: Right.
Speaker 2: Yeah.
Speaker 3: Yeah, so so the question is is um as we're moving forward, how do we sort of carry this work forward and sort of continue to address some of those legacy issues and those sort of quirks and strangestnesses of strangenesses of uh usage that accumulate over time. Um and and I I would just say like It's iterative, right? And one of the things we've actually prioritized, our design and development group is prioritizing for this year is something we're calling Wagtail democratization. So as I mentioned, like right now the use case generally for the users of our system is that Power users are basically the only ones using it to any great extent. It's like myself and people who are integrated into our development teams essentially using it and having a tight relationship. But obviously like Part of the purpose of a content management system is that anyone is supposed to be able, like the people who own the content, you, those are the people you want touching it.
Speaker 3: So we're making an active effort to make the system friendlier to those people. And what that looks like with a system with this level of complexity and like legacy culture is just iteration, iteration, iteration. So we start with what's going to make the most sense to our users right now? Then we have conversations with groups sort of progressively towards the people who are going to be the hardest to reach and say, what about this is scaring you? What about this doesn't make sense to you? And you just I don't know that there's a better way to do it than just to continue iterating, continue auditing, and continue having those conversations. Because every user group is going to be a little different, have their own anxieties and fears. And and something I have noticed happens a lot with systems like this is you get people who, similar to using Rogml because you just figured it's the easiest way to do it. You've been in the system long enough and you know that if you try to do thing X in situation Y, it's gonna fail for a reason that doesn't actually make any sense, but you've just encountered enough that you don't think about it.
Speaker 3: So you need to have conversations and bring fresh eyes into a system to be like What about this doesn't track to you because for your own team, it's the iteration needs to be coming from a place of constantly like, don't assume we've made the system as easy as it possibly can be. Um, because it's A, it's never going to be the easiest it possibly could be for anybody. But B because it encourages you to keep having those conversations with people, to keep iterating towards a CMS that manages the complexity without overwhelming a user who's walking in day one.
Speaker 2: Yeah, and I think too there's back to that, like there's some kindness that has to be involved both to your users and to yourself in that like you don't know now what your users are going to be doing with the system you build five years from now, right? So you just have to be responsive to the needs as they come. And again, don't question one of the most destructive things I think you can do is to look back at your own work and go, this is crap, right? And we do that all the time. I say that, but I do that all the time. But it's it's fine. And it you were competent at the point you made those decisions. You're competent now. It's just the context has changed.
Use Wagtail’s built-in page type usage report to see how many pages use each type, including page types used only once. Because page types are Django models, you can also inspect them with ordinary Django querysets.
Discussed at 10:59Walk the nested StreamField definitions recursively to collect every possible block path, then walk the stored StreamField values to find which paths are actually used. The speakers packaged this approach as a Python library and management command that exports block-usage data to CSV, with optional page-type filtering.
Discussed at 17:54They added proper Wagtail features and migrated the existing content: Draftail anchors replaced manually inserted anchors, icon and styling options were added to the relevant blocks, and other one-off HTML patterns were remediated. They then removed the old HTML-unescaping behavior that had allowed arbitrary HTML in rich text to render.
Discussed at 30:33They are usually trying to implement a needed pattern that the CMS does not support, such as anchors, icons, custom heading styling, or expandable content options. The raw HTML is evidence of unmet content needs rather than an attempt to break the CMS.
Discussed at 32:09Do not try to audit the entire site at once. Break the work into small, data-driven investigations, use reusable audit tools, and focus on patterns that can lead to concrete simplifications or better support for content users.
Discussed at 35:48Wagtail Inventory is useful when you already know which block or content pattern you want to locate and need its individual page examples. The custom audit works in the opposite direction: it exposes broad usage trends and all possible nested variations across a large site, without requiring a predefined search target.
Discussed at 37:42Combine an external crawler, which identifies the rendered URLs and visible patterns affected, with an audit of where those patterns live in Wagtail fields and blocks. This narrows the review from the whole site to specific content structures, making it possible to check and update duplicated or inconsistent versions.
Discussed at 40:35Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 19, 2024
Published July 19, 2024
Published July 19, 2024
Published July 19, 2024
Published July 19, 2024
Published July 19, 2024