New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Bluesky, atproto, how to hack on it and why it matters

ETHBerlinThu, Jun 19, 2025, 12:18 PM · 24:31

In 2019 Twitter asked a team of crypto and p2p devs to design a protocol for the next generation of permissionless social media. The protocol escaped Twitter and now drives Bluesky, the fast-growing social media site. It can do a lot more. This talk will cover: the principles underlying the design of atproto, why those principles matter, Bluesky and the unbundling of feeds, blocklists, curation and tagging, how the protocol works (Dag-cbor, Merkle Search Trees, DID:PLC), and how to make bots to make stuff happen on-chain.

Transcript

Okay, thanks Martin. So I'm going to talk about AppProto and BlueSky. I should be clear, I'm not going to be talking on behalf of BlueSky. I don't work for BlueSky, I wasn't involved in building AppProto. I'm going to talk about them because I like them.

What I will do is, if you go to BlueSky, to my profile goat.navy, I've put a bunch of links on there, which I will mention a couple of times in this talk. There you can find what the people who actually built AppProto say about it, which is going to be kind of a different angle to what I say about it. I don't think I'm going to outright contradict them, but maybe it's a slightly different angle. So, long, long ago, the internet looked kind of like this.

If I wanted to share something with my friends, I'd make my HTML page and upload it to my web server, and they would go to the web server with their browser and read my stuff. That wasn't a peer-to-peer system, right? It's a client-server system, but it has kind of a lot of properties that we're often looking for when we're building peer-to-peer stuff, right? So anybody can publish, nobody's stopping you publishing. You can self-host your web server if you want to, usually you don't bother.

But you don't really need to, because if you're using shared hosting, the hosting is unopinionated, right? The hosting has no editorial control. So if I just put my stuff on a regular web server and they start messing it around, I just move it to another web server, change the DNS, and there we go. So we have that nice property, but then we ended up with Web 2.0.

There were some good reasons why we moved to this different model, right? So we wanted to have our social connections. We wanted kind of identity built in. This tweet is funny because I know that this guy is the former spokesman for the Irish Republic Army. We wanted structured data, so your client needs to know that this is a tweet, and this is a like, and this is a retweet, and all that stuff.

And we needed resilience, because often, early on, if your content, if your stuff went viral and everybody hit your web server, it would fall over. So we needed that stuff fixed, and Web 2.0 fixed it by having like three web servers in the entire world, and they're always provisioned for loads and loads of traffic. So that gets us to a slightly different architecture here. We've got everybody's posts are on one guy's web server, in one guy's database, and everybody is going to that.

So that means that this guy now has editorial control over your life. He decides what people want to see, and we can all see the problems with that. So the lesson I want to take from that, which I will come back to a couple of times, is if you have a decentralized protocol, but it doesn't solve the whole of the user's problem, if you're missing some features, the market will put those features back with centralization. So one of the people who ended up in control of all our stuff was Jack Dorsey, a Web 2.0 entrepreneur.

Sometime in about 2018, he went from Myanmar and he meditated in a cave, and he came back and he read a thing by Mike Masnick, and he said, we should decentralize Twitter. I should not be in control of all your stuff. Twitter should be a protocol and not a product. So he brought together... They started I think with a mailing list, and then he ended up hiring Jay Graber, who at one point worked at Zcash.

And her job initially was to put together a team that would allow them to decentralize Twitter. So they built a system that was supposed to be able to... It had to be able to scale to Twitter, and it had to have kind of the user experience that you're used to on Twitter. So you couldn't say, OK, this is kind of cruddy UX, but it's decentralized, so it's OK. It had to kind of work in the way that people were used to.

So they designed AT Proto, or At Proto. I'm not even sure that's the right way to say it. They built the first application on At Proto, Blue Sky. And they ended up with a separate company, because this was very fortuitously that Jay insisted that rather than being like a project group inside Twitter, that they have their own separate organization that were kind of initially contracting to Twitter. So they had their own organization with the slogan, the company is the future adversary.

Right, so you know how Google had Don't Be Evil, which was nice for a bit. But you should just assume that if you're dealing with some organization, sooner or later the organization will screw you. And when that happens, what we want to be able to do with Blue Sky is not just be able to migrate off it, but ideally that process will actually be really seamless. So you should hardly even notice whether you're using Blue Sky infrastructure or some other infrastructure. So after this happened, Jack Dorsey went on Blue Sky.

So they launched Blue Sky and Jack Dorsey went on there and had a really bad time. And he stormed off in a massive huff. So he was initially on the board of the Blue Sky company and then he left the company. So now there's basically no former Twitter involvement in this fully independent project. So what does it look like?

Well, the absolute core of this thing, of the At Proto design is the PDS. And what I've done on this slide is I've rubbed out web server and I've written in PDS. But otherwise it's the same. So how is that different? For a start, the PDS is for structured data.

So the web server is like a document centered, but no, this is structured data identified according to a lexicon, which I will talk about in a sec. It's attached to an identity. So it belongs to somebody and it's going to be signed and I'll talk about. It contains in it the social graph of all the other identities of people you're referring to. So all of the entries in the PDS can refer to entries in other people's PDSs.

And the important part is it's content stressed and cryptographically signed. And that means that if you go to my PDS and get my content, you can then give it to other people and they won't have to trust you. So this is something that normally, well, often we would do traditionally for peer-to-peer systems because we're assuming that we're going to get our data from an untrusted peer. But here, we're not necessarily going to do that way. We may be doing it from a semi-trusted server.

But the important thing is that you don't have to get all of the content direct from the person who made it. And that also makes it really easy to scale because it means that I can run my PDS, I can telepost on a really crappy little VPS. And even if I get massive viral traffic from my goat pictures, my PDS is going to stay up because most people aren't actually going to be getting the data from my PDS. Okay, there's a bit more infrastructure in here. I'm still going to skate over most of the infrastructure and there's a really good piece by Paul Frazee in the links I'll give you at the end that gives you a lot more detail about the infrastructure.

But the core pieces we've got here is we've got the relay. And the job with the relay is to get all of the latest changes from everybody's different PDSs and bring them all together in a massive stream, what we call the firehose, which is just a stream of all of the changes happening in the entire world on everybody's PDSs. You then run a bunch of app views which are basically backend servers for the different clients. And the relays and app views, these are both things that Blue Sky are running them, but there are also a bunch of other people now also running their own infrastructure. So let me take you through what happens if I make a post.

So I'm going to make a post here that's referencing Neeraj's post, so referencing somebody else's post. So I've typed in my thing into my Blue Sky client. I've hit post. We're going to assign it an ID. This will be a PID, which is basically a time stamp with benefits.

And we're going to encode it in DAG-CBOR. If you're not familiar with that, DAG-CBOR is like a compact binary way of encoding structured data. It's designed to be very flexible, sort of easily round-tripled with JSON, and also kind of very compact and very easy to parse. Then we've got DAG-CBOR, which is kind of a more sort of subset of CBOR, which makes it kind of easy to parse. It does things like specifying the order of the keys in the document.

So we've put in my text. We've put in the type of text. And then at the end, we're going to hash that to make a content ID. So all of the identifiers that we're going to use to talk about our content are going to be hashes. This looks a lot like IPFS.

The CIDs look exactly like IPFS. And inside this one, where I've referenced Neeraj's skeet, I've got the URI to tell us where his PDF is, and the content ID so that you can verify exactly what should be in that quote. So I'm then going to add that to my personal Merkle tree, or my PDF. This is a slightly exotic kind of Merkle tree called a Merkle search tree. Again, in the links at the end, there'll be some explanation of why they chose a Merkle search tree.

But what's going to happen is that the content ID is going to be added in a new node in this tree, and doing that is going to update the root hash. We're then going to make the commit node at the top, so we put in there just our ID that we know about, and then the Merkle root of that tree. So then we're going to sign that, and then anybody who's got that commit node and the content and some intermediate nodes can then always verify that it was my content. Okay, let me talk about identity. So there are two standards for identity in AppProto.

There's did-web. This is an old W3C standard that predates AppProto. And then there's did-plc created by the AppProto team. I think it stands for a new thing now. It keeps changing.

What we're trying to do with did-plc is we want to get as close as we can to content address data, right? So ideally, what we'd be able to do is we'd be able to take the document that tells you everything you need to know about my identity and hash that, and that would get us to the identity. And that actually works for the first update. So the first time that you make an identity, we're going to put in stuff like the endpoint where my PDF is going to be. We're going to put in my username, which is the domain in AppProto.

We're going to put in the key that controls this identity. And we're going to put in the verification method, which is the key that's going to sign that Merkle tree whenever we make an update. So we're going to hash that, and for that first update, that's it, right? The ID is just the hash of this document. If we want to make a change to that, then we can.

We can go in and make basically the same document again, change whatever field we want to. So I might want to change my domain to my own domain, goat.navy. And then we're going to put in the hash of the previous entry in the prefield there, and we're going to sign it with one of the legal rotation keys. So if you've got a single chain of updates, you can then always verify that chain.

You can get the hash of the first one, and then you can check that all of the next entries are signed by the correct keys. What is slightly cursed about this is that you have the ability to change your rotation key. So if you change your rotation key and your original rotation key is compromised, then somebody could potentially make like a malicious chain of these updates, which nobody else would be able to tell without any additional information, which was the real one and which was added later. And Approto right now deals with this with a trusted sequencer. Like I say, you only really have this problem if you have a rotation key.

You could do it the Nostra way of never losing your original key. If you do that, you won't be subject to this issue. But this is the slightly kind of cursed thing that we're still dealing with. Okay, but it's a social network. A social network isn't just messages.

And remember I said early on, if you don't fully solve the problem, right, if you don't provide a way for people to do everything that they need in a decentralized way, then they're going to do it in a centralized way. And you will end up with, for example, one dominant client and everything packed into the client. So we've also got to do curation, moderation, stuff like that. So how does that work? Well, so curation.

If you go to Twitter, you've got the For You feed. This is something curated by Elon Musk. He decides what you're going to see. In Blue Sky, they have a thing called Discover, which performs the same role. It's not full of fascism, but it is rubbish.

But that's okay because anybody can make a feed. So I would normally use my Like of Likes feed. This is a really simple algorithm based on just like the people I like, what did they like? Some guy in Japan made this one. To make a feed, it's really, really easy.

So you listen to the fire hose, you watch all the changes coming in, and you pick out the IDs of things that you think are interesting to the user you're trying to send it to, and then you emit a bunch of IDs. And that's it. And any user can then just add that feed in their client and they'll be able to see it. Similar story for moderation. So Elon Musk on Twitter decides who is not allowed to be on Twitter.

On Blue Sky, the users can also decide who is not allowed to be in their feed. So you can do this in a positive way as well. So if you make a list of people, people can pull that in as a following list or they can use it as a block list. So I'm on the goat posters list because I post about goats. I'm on a technology list for people who are interested in technology.

And I'm on a tech shit list for people who hate technology and hate Web3 and hate everything about the whole thing, which to me, some days that's also me, but anyway. Okay, so everything I've said so far has been about Blue Sky. But Blue Sky is not the whole story, right? There's a whole lot more stuff we can do without Proto. So this is the key thing I want everyone to take away.

On this model, on the PDF model, everyone in the world gets a self-sovereign, cryptographically verifiable, semantic data store. And this is the correct way to do the internet, right? This is what we should have done when we went to Web 2.0. If you're making an application, you can write to that data store.

It's very simple. You can authenticate via OAuth with the user PDF, and that will give you the ability in your application to write records on the user's behalf to their data store. There's also a nice thing in the OAuth flow where you can put in, like, a PDF host and you can onboard the user into App Proto. So if they're not an existing App Proto user, you can still make your app work. So when you do this, you're going to want your data to follow some lexicon.

So there's going to be a data structure that you're going to follow. In BlueSky, you've got things like this, so a like has a record like this. If you're doing something else, almost definitely you want your own data, your own image of the world, so you'll make your own lexicon. So let me give you an example, a couple of examples of things people have built. Descent's like GitHub.

This is the same story, by the way, right? So Linus Torvalds makes Git. It's an incredible decentralized system. There's no need for any server anywhere. You can peer directly with the other people you're working with and you can always reconcile each other's changes.

But it turned out that wasn't all that the users needed, right? We didn't solve the whole problem with Git. People also wanted PR tracking and issues and all this other stuff. So they ended up using GitHub, a centralized server run by Microsoft. So we can avoid that state.

Things like Tangle SH is decentralized GitHub. All of the actions that you would be making on GitHub are instead made on your app protorepo. Then we've got SmokeSignal. This is an events management app. There's a really nice piece, which again I will link to by our friend Liggy, who probably many of you know, talking about the pain of being an events organizer and why we need something like this.

So the way I feel like we can use this, a lot of us have been building on blockchains. With blockchains, you've got some sort of commons data. You've got some data that's shared by everybody and you've got kind of a single history that everybody shares and has consensus over. And that brings you a lot of pain because you have to then limit who can write how much. There's a cost to everybody else to writing to it.

So you deal with gas and all these other headaches. But a lot of applications are not really needed. They don't really need a data commons exactly. What they need is data written by the user and controlled by the user. And when you're doing that, that's where I think you want to look at AppProto.

So all of these places, I don't want to be the huge hype thing, it's like it's bigger than blockchain. But I feel like there are more used cases in the Web 2.0 world that follow this model of user controlled history or should follow the model of user controlled history than that have consensus history. I mean, they're both important. Okay, I hope that makes sense.

So what I've done, I've put up a whole load of links on there. The Peter Frazee piece about turning the data center inside out is really great. There's some nice YouTube stuff. There's some other talks I've given, other events, and yeah, various other things. So please take a look.

Okay. Thank you. Thank you, Edwin. Thank you. Okay, we've got some questions here.

Feel free to keep asking questions or voting questions as we go through the Q&A here. We've got the first one here. Why are people in crypto not embracing AppProtocol more if it's value aligned in so many levels? In your opinion, is it in fact not value aligned or is it for other reasons? Good question.

I don't really know. I mean, I think part of it is, so crypto is centered around Twitter and a lot of people left Twitter because the owner embraced fascism, like American fascism, and started putting American fascist things into your feed whether you wanted it or not. And I think that large parts of the crypto industry like American fascism. So I don't think that's a problem. That's one thing.

Another thing is just kind of the way it's been vibed, that the initial blue sky user base was kind of, the defining characteristic was not liking Elon Musk. So there was also some kind of general sort of tech bro hostility which also applied to some of us. So maybe that's some of it. But I think we can do a lot more to make it more friendly to crypto people. Right.

Next one here is, what's the difference between ad proto, blue sky, and the Farcaster protocol? So I'm kind of having a hard time keeping track of the Farcaster protocol because they just changed it. It's a really different thing. So in Farcaster, the identities are on the blockchain, right? Which I think is very simple and works very well.

The posts are all shared by everybody and they're on something they make called like a snap chain, I think, which is currently a mission chain written by two people. And then the culture is very kind of money orientated. So that's the best I can say about that. Why not Mastodon? Why not Mastodon?

So I was on Mastodon. I was on Mastodon, but then the... So Mastodon... So, I mean, what's good about Mastodon, what's cool about Mastodon, is that there's no this will be finished later, right? Mastodon really is working with the proper community the way it's supposed to work.

The way it's supposed to work is, so blue sky kind of cuts horizontally everywhere, right? Blue sky unbundles everything and gives you freedom to take different bits of the stack. Mastodon kind of cuts vertically. So it says, basically the Web2 model is good, except that we need lots of different little Web2s with little different sort of fiefdoms and then they need to be able to talk to each other through message passing. So you're still then, because you don't have account portability in Mastodon, you're still kind of at the mercy of your administrator, right?

So I had a Mastodon account which I lost because the person running the server gave up. I think it was Lefteris. Somebody in the Ethereum community had an account and they got their account frozen for committing capitalism. So it's just a really different model. It's based on the model that you're going to still have this kind of Jack Dorsey figure, but you get to choose who they are.

And it's more a model like it really matters which server you choose on Mastodon, right? So the point of Blue Sky is the infrastructure should all be in the background. If you move, like, anything to do with curation, moderation, all of that stuff, you really want that to be something that you have an individual choice. So that's the difference in philosophies. Right.

How do relays discover the PDSs? I think the PDS is Pingdom, but I'm not sure. Anyone know? Implementation specific. Great.

Okay, we even got some helpful answers here from the audience. Are there other auth methods that can be supported for PDS access or discovery? How about delegation access to the same? Delegation access, not that I know of. No.

As far as I know, the only two ways for the existing PDSs are auth and app passwords, and the app passwords are delegated. Is monetization part of the protocol? Who pays for storage? So, as far as the, kind of conceptually, the person who pays for storage is whoever the user chooses to put their PDS on, right? So if I'm hosting my own PDS, then I pay my VPS host for storage, or if it's at home, then I pay Amazon for the disk.

In practice, right now, basically, most of the PDSs are on BlueSky, so BlueSky is like this massive PDS host that is paying for storage, presumed through venture funding or something. How do you know the data will be persisted on a PDS server? Well, it's your PDS server, so it's up to you to persist it. It's just your database. You can keep backups, and then you can, like, if it falls over, if you've got backups, you put it up somewhere else.

Is the BlueSky app completely open-source? Yes. That's my understanding. I haven't built it, but I'm pretty sure it is, yeah. Yes, it is.

And there are also a bunch of other nice open-source apps. By the way, have we done Farcaster? Oh, we did this different Farcaster. I forgot to slag off Farcaster about something. The overwhelmingly popular, basically the only really usable Farcaster application as of today is closed-source.

I don't know. Maybe we just end it there. We're at the end of the Q&A time anyway, so I think that's a great way to end it. Thanks so much, Edmund, for the talk, for teaching us about AdProto. Thank you very much.

Automatic transcript — names and jargon may be misspelled.