Activating Your Site: A Look at Activity Streams

This video features Ben Fonarov, Farhan Syed and Justin Quick at DjangoCon US 2014 in Portland, Oregon, USA.

Activating Your Site: A Look at Activity Streams
0:36:12
Published October 2, 2014
3,861 views

By, Justin Quick, Ben Fonarov, Farhan Syed
We will walk you through how to implement activity streams for your website in a generic fashion by leveraging the activitystrea.ms open specification. The two tools we will show are django-activity-streams which lets you interrelate the objects in your Django site using a supported DB and the activitystreams project which provides a ReST service based on the spec, with Neo4j graph for storage.

Help us caption & translate this video!

http://amara.org/v/FPWj/

Summary

Activity streams record actions by actors on objects and targets, then let users publish, follow, and consume those actions. They can increase engagement and provide data for recommendations, analytics, trends, and testing, but raise difficult choices around schemas, storage, scalability, celebrity and sparse users, query complexity, and real-time versus precomputed results. The presenters explain the Activity Streams specification and show two implementations: the Django Activity Streams package, which uses generic foreign keys, streams, template tags, and feeds, and National Geographic’s Horizon service, which uses APIs, Neo4j, Redis, optional Storm processing, and JavaScript clients. They stress that distributed implementations require careful handling of API chatter, caching, stale or unavailable external content, ranking, and data changes.

Key takeaways

  • Activity streams represent an actor, verb, object, target, timestamp, and descriptive metadata in a common structure.
  • Publishing actions and consuming friends’ actions can create a feedback loop that increases site engagement and supplies data for recommendations and analytics.
  • Django Activity Streams supports generic Django models, follows, built-in and custom streams, template rendering, and Atom or JSON feeds.
  • Generic foreign keys simplify flexible relationships but make aggregation and annotation difficult and can require inefficient or custom SQL.
  • National Geographic’s Horizon separates the activity service from content applications, storing graph relationships in Neo4j and retrieving current model data through APIs.
  • Real-time streams must account for scalability, graph direction, cache invalidation, ranking, external-service failures, and changing content.

Summarised automatically from the transcript.

Transcript

6,739 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:20

Speaker 1: Hey everybody, thanks for coming out. Um so without further ado, let's get talking about activity streams. Uh the talk today is activating your site, a look in activity streams. uh presented by us. We're going to go over briefly what they are, how you can implement them, and why you should care about putting them on your site. So just briefly, I'm Justin Ben Farhan. We are all work at National Geographic, which uses a lot of Django on some of our sites, and we've been doing this for a while. So again, here's the agenda. We're going to talk about activity streams, what they are, why you should care about them on your site, some of the uh engineering concerns around implementing them. And then we're going to talk through the open specification regarding activity streams and then

1:05

Speaker 1: two solutions, one expressly in Django, and the other one as a service that can work for any website. So what are activity streams? So they're actually everywhere. So GitHub, Facebook, LinkedIn, Twitter, etc. Um, you probably run one of these at least once today. Um They essentially are a way of displaying uh actions that people take on your site. Um and uh they also have a way that people can somehow like socially link themselves to other people who are also making activities and amplify the amount that they're seeing in their stream. So why do all this stuff? What's the point to all this? And it actually turns out that it's really good at increasing engagement on your site. So even if something as simple as putting a like button on your site will drive track traffic immensely.

1:51

Speaker 1: Because essentially people like clicking on things. Uh we try this out on NatGeo and uh just very basic, without getting any a lot of like user feedback or anything like that, people just like clicking things to say yes, they approve this content. or whatever. And also, so that's the first half of being able to publish your your activities to a site. The other half is being able to consume them and see what your friends are doing. And with those two in combination, you actually set up a really powerful positive feedback loop that will drive lots and lots of more content and engagement on your site. And that's a really great thing to do. Um but in terms of engineering, we get a lot of data out of it. Hooray data. And with that data, we do a lot of fun stuff, including tons of analytics. You can drive recommendations like like Netflix, show you related content.

2:36

Speaker 1: We have social graph apps. We can show you trending, can drive A-B testing, and tons more. Um but there are some uh engineering considerations that have to be taken into account implementing this. And uh Ben's gonna dive into that a bit more.

2:52

Speaker 2: Hi. So um these are just some of the problems that are presented with activity streams. Uh there are no great solutions to this. It's not a simple, clear-cut answer. But it's all about weighing the benefits of each one of them. I'm going to dive into a little um into one of these into each one of these a little bit. So the first problem is too many peppers. Essentially, there are so many implementations of activity streams out there, and that's kind of evil We don't really get the benefit of cross-site implementation. It would be really great if I could do an activity on Facebook or somewhere else and then get that same thing at National Geographic or another site. It causes duplication of work and there's no common semantic structuring, so it's kind of hard to implement these things and come up with the terms on your own every time.

3:39

Speaker 2: We're going to talk a little bit about the solution to that in a little bit So the second problem that we might face is what do we actually store and what type of data structure? And in what schema? Is this going to be a relational database, a non-relational database? What's the data structure? There's different things in terms of efficiency that could come into this. Are we doing an adjacency list, an adjacency matrix, uh doubly linked list, something else, a key value store? just a hash. Are you going to have more writes or reads? So if you're a social media company, you might show the newsfeed a lot and then you 're going to have a lot of reads. If you are a company like National Geographic, you know, there's mostly content. I'm not going to show you the newsfeed that much, and I'm going to have a lot more rights. And then do I store the activities in the same place that I store the actual stream?

4:27

Speaker 2: Where was it pre-computed or something like that? I'm going to touch on that in a little second. So another problem is centrality versus sparsity. Haters gonna hate. You're always gonna have your celebrities. So those are people who have a lot of followers or do a ton of activities. And the reverse to that, which is sparsity, you know, forever alone. I only have one friend or I've only done one activity. And each of these things present a different problem. uh for instance for the celebrity do you start to enforce limits at some point that's going to be too much to compute um Facebook if I remember correctly has a limit for 5,000 friends And do you solve that with a data limit? Is that a hard limit? Or do you do some sort of UX solution where you hide the follow button at some point? And then the other way, uh you know the other problem is

5:15

Speaker 2: what do I show a person that has one friend? What do I show a person that's only done one activity? Am I going to show them stream with that one thing? Um A stream that has only that one friend's activities. Scalability, obviously an important factor here. You know what am I what am I dealing with? Am I dealing with a massive amount of activities or just a few? Where do I do this computation? Am I offloading it to somewhere else? Am I doing that inside the request-response cycle or doing something else with it? And how do I handle complex queries like friends of friends or follows or recommendations? And then uh real time versus pre-computed. When do you do your computation is an important factor. Are you pushing out in real time? If so, how do things take precedence? And are you doing some sort of half and

6:00

Speaker 2: half? where you're pre-computing some stuff and then computing on the fly for other things. And what are you actually sending? Are you just calculating the ranking of things? Are you doing the entire thing including the HTML and then just sending that off What are you really going to compute? So I'd like to invite Justin back on, and we're going to talk a little bit about the solutions to these things and some of the implementations that we came up with.

6:27

Speaker 1: So there is a solution for some of these problems. Um it turns out there's an open uh specification called the Django, the activity stream specification. uh that aims to solve at least you know the varied implementations by solidifying everybody on one uh semantic structure and then the data and how it's stored. And so this has actually been around for a while, and all these companies uh here are implementing it in some fashion or another. Uh we implement this at National Geographic and the other solutions we show you today. also implemented. It supports uh Adam and JSON uh data structures out of the box um and uh in a a human-friendly and machine processable way Quick disclaimer, nobody up here is officially involved with the specification. We're just implementers who are showing our um sharing our experience with you about this.

7:15

Speaker 1: So the uh spec looks a little bit like this. So there are three big entities. Um the actor, which is the required one, normally that's the user on your site, is whatever's taking that particular action. The verb phrase, which is uh whatever uh actually happens like commented, post, um, the spec defines a whole list of officially supported verbs. But for sites, most of the sites I've seen this implemented, they take those and then the ones that aren't there, they sort of just run with and customize their own. There's an official channel to go back and propose verbs that aren't there to the draft specification. Which I encourage you to do on their website. You also have an action object, which is the primary object of the activity, whatever gets created essentially, whether it's a photo in an album or a comment on a blog post.

8:00

Speaker 1: And then last entity is the target, is wherever that action is taking place. You know, it could be in like like the album or the blog post, wherever the action is going. And then there's a little bit of meta information like, well, what the timestamp is, uh, and a descriptive title and summary And so this is what the JSON uh implementation looks like, uh really quickly, uh quick example. Uh you see that the actor object and target have their own uh JSON object right in there. Um And we sort of stuck to sort of like Jason as the first class citizen for the rest of our implementations. So I'm going to talk to you right now about Django Activity Streams, which is a uh an open source project that I wrote. uh and has been used in several different sites. So it uses the specification. It can track any uh object in your Django

8:47

Speaker 1: project. Uh it runs on any supported Django database uh and it keeps track of everything using generic foreign keys. It also provides a way for you to render these streams onto your site using template tags or feeds. Everything's generated at request time, and I leave the caching up to you. And you can read along with the source code at my GitHub repository or also read the docs. So When I'm not working at uh National Geographic, I'm CTO of Narwhal Studios, which is the gaming company that uh runs the humans versus zombies game. And HVZ is essentially a uh organized game of tag. Let's play it. colleges, universities, and other locations all over the world. And uh we deal with like thousands of players doing hundreds and hundreds of actions a day.

9:33

Speaker 1: And uh we needed a way to sort of give that feedback and show people their streams. Um and so this is HBG Source, the uh Django site that we set up to help moderate that game. And you can see a simple action right here is just the row. Crazy face the actor uh selected the verb. Uh the original zombie is the object. And the demo game is the target, and then the timestamp is in the time since filter, all displayed right there. Um so behind the scenes it uses two models to accomplish everything. The main action model has generic foreign keys named actor, target, and action object that can point to any object in your Django database. It doesn't have to be a user. It also has that descriptive meta information I was talking about as well. The second model is follow.

10:19

Speaker 1: It has a foreign key to your user, whether that's Django Auth user or it also supports a custom user. That's up to you. And then a generic foreign key you point into any other entity that you'd like to follow. It doesn't have to be user, it could be anything else. And uh it maintains that relationship in the database. So uh generating actions are pretty simple. You just import the uh signal and send it along with the arguments. So crazy face is the actor first and foremost, and then all the other arguments are sent in as keyword arguments. So Crazy Face selected the original zombie for the target, which is demo game. Following and unfollowing is similarly also simple. There's just a follow and unfollow function that take a user first and then an entity second and either create or destroy that

11:08

Speaker 1: follow object And there's also followers and following, which would return a query set of users that follow that given entity, or the other one does the reverse lookup and gives you a list of entities that that user is following. And of course there are streams. What good is this app if you can't show it? So it comes with a few built-in streams. Um user stream being the most important one that takes a user. finds their your their followers and then gives you actions that those followers have done. That's like your main dashboard of GitHub or Facebook or Twitter. That's probably the most important one. And then there's similar other ones that are actor, target, and action object that just do a similar lookup based whatever context that that given object needs. is in and that action.

11:54

Speaker 1: So uh actor streaming crazy face will show me all the actions where Crazy Face was the actor. Uh model stream is another interesting one. It'll show you any and all content or any and all actions that involve a particular content type. So this will show me anything that is happening with any user model. And so if I were to graph out the interesting one, the uh user stream query, essentially you're given the uh the user object, it goes and finds your followers. or the sorry, the p the content that you follow, and then it reverse looks up uh the actions that were um involve those objects. Uh and it can also filter it down by relationship. Uh by default, the user stream just returns everything wherever your uh follower was involved in, but you can customize that as well.

12:39

Speaker 1: Um so those only get you so far. Uh there 's also uh custom streams which are easy to implement. Uh so basically uh this first one right here. uh player actions takes a game instance, uh it finds out the number of player IDs that are in this game, and then it returns a query saying, uh show me all the actions where the players of this game were the actor. And it does that by object ID and content type. The stream decorator gets you a little power where uh it will you can just return query set arguments or keyword argument filter. Um you don't have to return a query set. And then the second example here is the player actions by slug. It simply does the same thing that the first one was doing, only it takes a slug instead of an instance. We're going to get back to that guy a bit later about why that's useful.

13:25

Speaker 1: So to graph that out, it looks pretty similar to the first example. You're given a game, you find the players in that game, and then you do a reverse lookup to find out where those players made specific actions when they were the actor. So it's a little it's it's pretty straightforward. It's the generic foreign keys work from the uh player objects to the actions uh and keep track of the relationship that way. So the first way to put this on your site is uh using template tags. And the activity stream uh template tag is the most helpful because that essentially takes The first argument is the name of the stream you're interested in, like the actor stream. Then any arbitrary objects you want to pass in through that tag. This returns a stream object in context, which you can iterate over and display however you want. There's a helper built-in called display action, which is just an include tag

14:14

Speaker 1: that renders a specific template that you can override in your project, but you can also display this however you want on your site. It also works for custom streams, so you just give it the name of your custom stream, pass in a game instance, and you're golden. The second sort of overall function that the templates give you is ability to create follow uh buttons. So this guy essentially is a uh a link that's a toggle. If you're not following this person, it'll give you a link to follow them. If you are following this person It will give you a link to unfollow them. It's sort of like that toggle that you see on GitHub or any of the other uh social media sites. The second way to get uh information out of the app is through feeds. Um so the user feeds first and foremost. Um Any of these support either Atom or JSON

14:59

Speaker 1: for the user feed, as long as you're authenticated and go to that URL, it will return out the user feed in the machine readable format Object feed does a lookup for a specific object by content type and object ID and then returns you a stream of actions where they participated in that in any any relationship. Model feed just like I showed you just does uh everything based on the content type. Interestingly, you can also add uh custom JSON feeds like this, like I showed you with the game slug, um the second custom one before You can actually pass URL parameters directly into your streams and render things out through a custom JSON feed that way. So when implementing this, I ran into a couple uh database considerations that are important to note. So um if you were to write this thing uh

15:46

Speaker 1: naively uh and write this out on a template, uh you'd essentially get one query for getting all your actions. And then as you were iterating, you would do a hit to the database for every actor, every target, and every action object. And this gets incredibly complex after a while. So luckily in versions of Django 1. 4 newer, you have prefetch related, which essentially combines it all and it just has big O of C where C is the number of content types overall. So it drastically reduces the number of database queries. It makes the queries a bit beefier. And then Django gets that information back from the database and then shuffles everything together in Python to give you your final query set. So Django Activity Streams use this under the hood. You don't have to worry about this, but you can extend it if you'd like

16:31

Speaker 1: uh further on. Also since we're dealing with generic foreign keys, there are some limitations. Um like the aggregation and annotation API of Django will not work for generic foreign keys that I found. So like this guy right down here that tries to find account of actors whose health is greater than the five just will will not work. So this leaves unfortunately this leaves out a lot of the interesting things like um Recommended content, like uh most popular, like a lot of the more interesting queries that you'd like to get a handle on, it sort of falls short, and you have to do a lot of ugly SQL to get things the way you want to do. So there's a sort of a better way to do it, and that's what we've been using at National Geographic. We've been uh working on the horizon service. Uh

17:16

Speaker 1: and I'm going to toss this back to Ben to tell you a bit more about it

17:20

Speaker 2: Hello again. So a little bit of background about the Horizon service. Like Justin said, if you're building an app which just needs an activity stream, His project is awesome for that. We have a large ecosystem. We don't really have control over the models that exist within the different sites that we have. And so we needed some solution that was able to deal with this type of stuff as a service. We have a great product owner that basically told us there are three different implementations of favoriting right now. Build me one. And so we built a service. And it follows the activity stream spec or tries to. There's some limitations with uh some of the solutions that we chose. It does a half and half sort of real-time versus pre-compute mix And

18:06

Speaker 2: what's important to know is that if you want to use this, your models must have an API. We'll talk more about why in a minute, but That is a requirement. And then there's clear separation between front-end and back-end modules that come together with this. And this again is open source. So, a little electrical circuit for you about the Horizon ecosystem. I'm going to dive a little bit into those and we'll talk about them. But first, storage considerations. Uh first thing we look at was what do we store this in? And a graph database really is perfectly suited for these types of things. And for making interesting queries. It's optimized for large traversals. It's really good at storing the relationships and we can look at individual slices based on different things, which I'll talk a little bit about

18:55

Speaker 2: in terms of the choices that we made. So Neo4j was our um product that we chose to use for this, the database. We looked at others, but eventually chose Neo4j. It's based on Tinkerpop, which is a Java framework for property graphs. Titan, which is another graph database, also uses the same thing. Property graph is essentially a graph database that allows you to have properties on both nodes and edges. And underneath the hood, it implements a doubly linked list as its data structure for relationships, and nodes are just pointers to their first relationship, and that's how they traverse. In terms of complexity, indexing and search are a little bit costly because it's a doubly linked list. It has a big O of N, where n is the number of edges. But and then for insert and delete, it's constant speed.

19:42

Speaker 2: But Neo4j actually uses Lucene on top of things, so when you start when you start talking about indexing and search, it gets a much better result. But what do we actually store? So we don't want to store your entire model. We can't actually store your entire model. Our situation is one where There's multiple models we could have a photo in one site described completely different from a photo in another site. And in trying to solve that and in using the bay the the best practices for something like Neo4J, we decided on five basic properties for nodes. Those are API. If you remember I talked about you needing an API route for your models. An AID, your application ID. a type of app label model name, so let's say YouTube underscore video, created and updated timestamps, and then that's

20:30

Speaker 2: That could be an actor, an object, a target, or anything else in your graph. That's the only thing that we store. For edges, it's a little bit different. Name4j has a native type. So that could be followed, favorited, um, liked, watched, so on, and then a created and updated timestamp. We also use Redis. We use Redis for sockets, sessions, caching, and then some stream data. It's really good for that type of stuff. There's not much to say except that it's an excellent database. And then uh back to this part. I'm going to talk a little bit about the pre-compute cycle. So Um we said that you have to make a choice between real-time and pre-compute. We said that we chose half and half. We use

21:16

Speaker 2: Apache Storm. I don't know if you're you're familiar with it. It's a really great product. It's like Hadoop, but for real-time message processing. It's not a dependency but a recommendation. And the way that Storm works is it has these topologies. Essentially topologies or Storm in general is a processing framework that is distributed and really good at doing these types of things for processing messages. It has topologies that describe a set of processes. Those are called bolts. And you could have multiple topologies. You upload a topology by just uploading a jar file into the Storm cluster, and then it runs those things. It internally can run not only Java but Python, uh JavaScript, C sharp, whatever you want. And communication to it is done via message queue. Um we use Kafka or RabbitMQ

22:03

Speaker 2: for some certain things Storm topologies are really good because we had a problem to solve, which is one part of our company might want to have larger weights on videos, and another part of our company might want to have larger weights on articles. And in order to do that type of calculation and computation, we can create many different storm topologies that basically define different processes to eventually get us the data that we want for each one of these storms. streams. All of that is dumped eventually into Redis. And I think I said this, but I'll say it again. This is not a dependency, but a recommendation. You could do everything on the fly for the horizon ecosystem. So Horizon itself is built on Node and Sales, which is an MVC framework. And I'm going to dive a little bit into the API for it.

22:48

Speaker 2: So it supports multiple content models with the use of the simple namespacing, which is the app label model name, and allows uh access to activities from different viewpoints. You could look at things from the actor viewpoint From an object viewpoint, that'll make a little bit more sense in a second. Um, and then uh target and so on. Here's an example call for you. So we're at version one of the API and if you go to object, YouTube Video 1 and then Favorited, essentially you're going to look at every activity that is of the type favorited that has been done on YouTube Video 1. And we're looking at that from the direction of the object. So I'm asking what has been done to me, the object. And on the right you can see that we uh follow the spec and you see the little data parameter there.

23:35

Speaker 2: That's actually an after-effect of Neo4j. uh which we're working to overcome, but uh so we try and follow the spec to some degree. Um here 's some more examples. If I did actor auth user one favored YouTube video, and then if you go to that YouTube video, you'll find something funny Um that would return a specific activity as described by the spec. And if you went to the object YouTube video with that ID, favorite auth user, I would see all the activities done to an object by a specific type of user. And that's useful for accounts, for instance. Thing to remember here is that direction actually matters when dealing with graph databases. So if I look at things from the point of view of an actor, it's not the same as looking at it from the point of view of an object.

24:21

Speaker 2: I'm asking different questions. And actually within the graph databases, sorry, within the graph database, edges have direction. So if I'm asking, I'm an object, what have I done? I'll probably get nothing because most YouTube videos can't like things. Uh but if I'm an actor and I'm looking at from that direction, then I'll be able to get some results. Posting, really easy. API V1 activity. Basically the payload looks More or less like the data that we store. Um and then we do manipulations on top of that for the pre-compute stuff And you could even do complex stuff. So we created a controller that's uh called proxy control. There's also a reverse proxy controller. This facilitates stuff like follow. So in this graph example, you have a proxy that could be you, and let's say I've

25:07

Speaker 2: the proxy verb that I do is followed And whoever I follow, that could be actors or objects or whatever, what I'm going to get back is a list of all the activities that those actors have done on objects. The return call would look pretty much the same. The return result would look pretty much the same as what I showed you earlier, except that in this case I would have multiple actors doing the activity. So what are the problems that we run into uh something like this? First of all, we have no control over external data. Um if uh If a you if a photo changed its title, I don't know about it. So you really need to live in an ecosystem that allows you to get that data back and to inform you of such changes. Second problem that we have is that graph databases don't really have a great ecosystem yet.

25:54

Speaker 2: So the adapters are not that great, there's no real graph ORM, and I know because I wrote some of those adapters And then front-end versus back-end computation. I heard that I missed a great talk this morning about Where to do the computation and that's a real consideration. Do you do the entire stream processing on the back end? Do you send it up to the front end to do uh to do some of the computation? And that leads me actually to the next part uh which It's going to be Farhan. Farhan uh worked on every part of this uh ecosystem, but he's going to talk about the front-end modules and how they relate to this ecosystem. Thanks, Ben.

26:39

Speaker 3: So what are the client-side modules? Uh so you see there's the stream and snippet. Um that's just the names we call them, and they're just standard web front-end technologies, HTML, JavaScript, CSS. Um and they allow you to communicate with the Horizon service about by sending actions and by consuming actions. So let's dive in. So the first thing we're going to talk about is the snippet. It's like a like button. It's a very configurable like button. The snippet is responsible for representing a specific verb that an actor can take on a specific object. And it's also responsible for displaying some state about the activity of a specific object, like counts. So here's some representations of those. And

27:24

Speaker 3: it's also an open source process, please check it out. And this entire thing is built with just vanilla. js. So say I have a web application and say I have an awesome picture of me and Lamfort Burn. And I want other users to tell me, hey, if I like this, tell let me know if you like this picture. So on my blog or my web application, I can just add this div and the snippet will be more or less appear And you can see the div is pretty standard. We're using standard HTML data attributes. There's an object type, and there's again the app label underscore model name. The AID, which is the application ID, so wherever whichever application has the storing this photo, and then the object endpoint We'll come back to that later, that's really important.

28:11

Speaker 3: And then the data verb. And as you can see, a snippet is kind of a map to an object. Um and that and the snippet and the stream have this concept of context. Um the snippet since it represents an action you can take on an object, you need to ask the question, well who's taking the action? And a lot of times it's gonna be a user So most likely it'll be a user, but it could be really anything. And we mentioned it was highly configurable. So the dataverb attribute actually maps to a template. And so you can kind of really easily customize how a verb looks differently and how and even the business logic that encapsulates it. And it's really easy to add create your own verbs at the very bottom. Say your application needed a new verb, Pipered, you know, whatever that means for your application.

28:56

Speaker 3: You can very easily add an attribute, make sure you have a template that's associated with it, and then you will get a custom verb. So let's look at kind of how this all works and how they are kind of look like the request-response cycle. Um so let's say I have the snippet, it has counts of eight, and the user clicks the snippet. This sends a post request to the same endpoint, and this is basically the payload. Again, it's pretty exact we try to be really consistent. So it's AID, API, type, app name, model name. The Horizon service returns an OK, the snippet is updated with the new state, the heart is filled in, and the count is all there. Um so I just showed you one example of like one snippet on a page, and it's like, hey, can there be multiple snippets on the page?

29:44

Speaker 3: Can multiple snippets be pointing to the same object?

29:47

Speaker 2: Yes, they can.

29:48

Speaker 3: We kind of solve a lot of these issues. Because of the variety of national geographics that pip you know, we didn't design all the all the front end web pages, so we need to have a way where these can all talk really easily to each other and communicate. So yes, one object can be represented by multiple snippets, multiple snippets can represent multiple objects on the same page, and they all work fine. So that's pretty much the snippet. We're gonna get talk about the stream now. So if the snippet was actions you can take, the stream is consuming. It's like it's the newsfeed. It's it's a highly configurable news feed It displays activities based on an actor. Unlike the snippet which is built in vanilla, this is built with just backbone. And again, it's an open source rebuild.

30:34

Speaker 3: So let's show you what the stream looks like. Right now you're viewing the stream of Lucas Servan, who's one of the developers at National Geographic, and you can see kind of the structure of the stream. Lucas Servan, the actor Favorited the verb, the article, digging Utah's dinosaurs. And on NGM is like the target in that case on some application. He liked a lot of things on July 2nd. I'm not sure why, but I think he was in a good mood. So let's go back to the example of this makeup, this is my my web application. So let's say some user has favored a photo and now the stream is showing you what they've done. And this stream is configured to show all of this particular user's activities. He favored a photo. And the most popular activity, which is Django

31:21

Speaker 3: ate some some child. Um let's get into the response request cycle here because it's really there's a lot it's a bit more complicated than just posting or deleting sending a delete uh request. So the module uh actually uses a WebSocket connection with the horizon service. Um the module because of this we have like a bi-directional communication So the module can ask, hey, give me all of my activities, or give me all the activities of people I follow, and what you're going to get back is a payload from the Horizon service. Now this is where the API endpoints really come into play. Again, we don't store any, we're not storing images, we're not storing model data, we're actually just storing some API endpoint that we will call out. So the Verizon service, I mean the module.

32:07

Speaker 3: will call out to all these external applications that your models are on and ask for them, hey, describe this image. So it's really important now to see like why you need these API. And this is why this these snippets in this module can kind of live in multiple web applications. They don't really need to know anything else. They just communicate through this mechanism. And when we get this response back, um we cach we cache all this in local storage. So you don't have to be constantly making these, you know, if you reload the page modules on, you're not gonna have to constantly make all these calls out to these external services. We cache all that locally. So this actually brings up a ton of problems and interesting considerations that we want to stress if you want to start using any of these

32:52

Speaker 3: One of them is that the API obviously is really chatty. The module is really chatty. You need you're supporting an ecosystem where there's multiple web applications. So be aware of that. You're going to be making lots of calls out. you might be making some course calls out, so just be aware of architecting that out. Um the and because we're relying on all these external applications, how do you deal with failure? What do you do when the service this other application fails? What do you display on the stream? Do you ca do you do you display some cache results? You know, these are all really important considerations to think about. And I mean there's I'm just gonna go over a few more. So Ben had mentioned this, you know, the service is not really aware of changes on content. If some photo gets updated here or some video 's title gets changed. How do you reflect that back to the service?

33:39

Speaker 3: And if you're already displaying that on the stream, you know what do you do? How do you how do you change up? How do you invalidate the cache? And a really simple example, even like sorting and kind of a ranking. Let's say I want to display the most recent popular activities on a stream and someone generates a new activity. Where does that come out? How do you do you immediately put that at the top? But wait, you have you're using an algorithm that's like, hey, I want to only show the most popular one. So how do you kind of navigate this realm? And then again, how do you a lot of this has to do with how do you invalidate the cache? Basically the point is there's a lot of business logic between the serv the service and these front-end modules. And so what you really need to do is determine what kind of updates, models, and things you care about

34:25

Speaker 3: so you can display them to the user. The good thing is that we've built the Horizon service in a way where the backend, the service itself, and the client-side modules are really configurable and so you have a lot of opportunity to kind of like tweak what you need. Is this running live? Yes, it is running live. You can go check out ngmbeta. com. Um ngm is the online version of the magazine where you can see all these issues going back to like 1888. So we encourage you to go on, become a member, and start uh clicking on and favoriting uh activities. Um This is what the stream looks like. This isn't released yet. They're still kind of working on this, but eventually this will be displayed on s on the user profile page.

35:10

Speaker 3: So that's pretty much the agenda. Oops. Yeah, that's pretty much the agenda And uh one more thing to note that all the projects we've spoken about, all four of them, all four four repos, are open source. We've definitely taken contribution uh contributes There's a lot we need to do and we're also building out a Django Horizon app that allows Django to speak because we have so many Django applications in Nat Geo. So Um that's it. And if you have any questions and we'll I guess we'll supply the answers, but let's try.

Questions this talk answers

What is an activity stream?

An activity stream displays actions people take on a site and lets users socially connect to and amplify one another’s activities. Common examples include the feeds on GitHub, Facebook, LinkedIn, and Twitter.

Discussed at 1:05

Why should I add an activity stream to my website?

Activity streams can increase engagement by encouraging users to interact with content and see what their friends are doing. The resulting activity data can also support analytics, recommendations, social-graph features, trending content, and A/B testing.

Discussed at 1:51

What are the main engineering challenges when building activity streams?

Key decisions include choosing a data model and storage system, handling high-activity users and sparse networks, scaling computation, supporting complex relationship queries, and deciding what to compute in real time versus in advance.

Discussed at 2:52

How is the Activity Streams specification structured?

An activity consists primarily of an actor, a verb describing what happened, an object involved in the action, and a target where the action occurred. It can also include metadata such as a timestamp, title, and summary, and the specification supports JSON and Atom representations.

Discussed at 7:15

How do I implement activity streams in Django?

The open-source Django Activity Streams project tracks arbitrary Django objects with generic foreign keys, supports follows and unfollows, and provides user, actor, target, object, model, and custom streams. Streams can be rendered with template tags or exposed through Atom and JSON feeds.

Discussed at 8:47

How can I avoid excessive database queries when rendering Django activity streams?

Naively iterating over activities can issue separate queries for every actor, target, and object. Django’s `prefetch_related`—used by Django Activity Streams—reduces this substantially by prefetching related data and combining it in Python.

Discussed at 15:46

What is the Horizon activity-stream service, and what does it require?

Horizon is a service-based activity-stream system intended for organizations with multiple applications and models they do not centrally control. The participating models must expose APIs, and Horizon separates the front end from the back end while combining real-time processing with precomputation.

Discussed at 17:20

Why does Horizon use a graph database, and what does it store?

A graph database is well suited to activity streams because it stores relationships and supports large traversals and relationship-based queries. Horizon uses Neo4j and stores compact node properties such as an API endpoint, application ID, type, and timestamps, rather than entire application models.

Discussed at 18:55

How do the Horizon client-side activity-stream modules work?

The snippet module represents an action such as liking or favoriting an object and can display its state and counts. The stream module acts as a configurable news feed, communicating with Horizon over WebSockets and fetching the underlying content from the APIs of the external applications.

Discussed at 26:39

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon US