How to Decentralize Any Front-End
ETHBerlin·Thu, Jun 19, 2025, 08:06 PM · 35:00
This workshop is about decentralized storage, one of the key components to realize the World Computer. We'll use the time to run a storage node (Swarm), upload a website to the decentralized network, and set up ENS to make the site accessible also on web2 via eth.limo. As we go through this demonstration, we'll also discuss challenges and solutions in decentralized storage, such as data availability, DDoS and censorship resistance, erasure coding, incentive systems, mutable vs. immutable data, and more.
Transcript
Okay. We are starting. Hi, everyone, and thank you for being here. This is the how to decentralize any frontend workshop. It's going to have two parts.
In the first one, we talk about decentralized storages in general, how they are built up. And in the second part, we do a quick demo where we upload content on this forum network and set it up with DNS so it's accessible to easily more. So, where do we start? To understand how decentralized storages work, I found that proximity order is pretty much your starting point. So, what's proximity order?
Proximity order is a function that takes two parameters, two addresses, and tells you how similar those two addresses are, how close they are to each other. Okay? The way it does that is that it counts the number of matching significant bits in the two addresses. Significant bits usually means leading bits. I'm going to give an example here.
PO, which is short for proximity order, I'm going to call that function with two addresses. Normally, addresses are 20 bytes or 32 bytes. In this case, it's going to be just dummy addresses which are two bytes. Okay? So, the first one is, let's say, C1D0.
The other one is C0D0. Good. And we get 7 for the proximity. Okay? So, here's the binary representation of the first address and here's the second one.
And you can see that they match up until bit 9, so they have 7 matching bits. Of course, if I pass in the same address twice, I should get maximum proximity, which for 2 bytes is 16 bits. And if I pass in two addresses which do not share a leading bit, I should be getting 0. Okay. So, in our decentralized storage, we have two very important base entities.
First one is the chunk. Chunk is going to be our unit of storage. Now, every chunk has an address. Our other entity is the node. Node is the software that you run, that you use to interact with the network that connects to the other nodes and so on.
Also, they are responsible for storing these chunks, uploading and downloading them, and syncing them with each other. We are going to denote the network with this uppercase N. Okay. So, we have the nodes and we have the chunks. So, let's start defining rules for this decentralized storage network.
Okay. So, the first question is, where do chunks get stored? And we can use proximity to give a solution to this. So, we are going to define a rule, this first one, that the destination node or the storer node is the node which is the closest to the chunk address. Okay.
The node which maximizes the proximity order with the chunk address. Good. We know which node is storing a chunk, but this immediately has an issue. What happens if that node shuts down and leaves the network? So, in that case, you have data loss.
Okay. So, storing chunks on only one node is not ideal. So, we are going to store them with a group of nodes and that group is called a neighborhood. These neighborhoods form based on proximity, nodes proximity to each other. Okay.
So, we are going to use the value D, network depth, and we are going to visualize it. So, here's our network. We have 10 nodes, so the yellow dots are the nodes, and they are at depth 2. That means neighborhoods are forming, there are two matching leading bits in the node addresses. Okay.
So, in this first neighborhood, all three nodes have binary 00 bits for their node addresses. Okay. So, let's say that something bad happens and half the network disappears like this. Now, we can see that we have three neighborhoods at risk because only one node is storing the chunks for that neighborhood, and if they were to leave the network, we would again have data loss. So, what the network can do and should do is that in situations like these where neighborhoods are underpopulated, the network depth should be set back to 1.
Okay. Now, only one matching bit needs to form a neighborhood, which means we have now, again, sufficient redundancy, but of course, this comes at a cost. Okay. So, now, one node has to store more chunks, has to do more work. The overall capacity of our decentralized storage network is less.
So, ideally, of course, you want a lot of nodes in your network. Okay. So, let's give our network a lot of nodes. We see that we have so many nodes that it doesn't even fit on the screen, so that means we can safely set back network depth to 2. We still have enough nodes.
We can set network depth to 3 and maybe even try 4. That means that now we have 16 neighborhoods. That means more total capacity for the network, and even with so many neighborhoods, we still have sufficient redundancy for every neighborhood. Okay. So, that was the core idea.
The rest is pretty much just implementation details to build the whole network. Now, naturally, there are two network types. The first one is the altruistic network. Those are free to use. If you want to make sure that your data stays alive in an altruistic network, you will need to participate in that network and, for example, pin your content locally.
The other type of network is the incentivized network. So, these networks usually have a token. You need to spend this token to buy storage, but this type of network does not require participation. So, you can get the token, run your node, upload something, disappear, go on a vacation for a year, come back, and you should still be able to access your data. This is because other nodes are incentivized to be in the network and store data.
We are going to explore an incentivized network, and for the incentivization part, I'm going to introduce a concept called the postage batches. Okay. So, so far, we haven't done anything on chain. This is the first time we are going on chain. So, to incentivize the network, we are going to use postage batches.
These postage batches are created on a chain by spending tokens. What's interesting about postage batches is that they have 65,000 buckets. This is a constant in network. These buckets range from, there are two bytes, and they range from all zeros to all Fs. And what's interesting about buckets is that they have some amount of slots.
Okay. So, you will have the postage batch, postage batch, each postage batch has 65,000 buckets, each bucket has some amount of slots. The amount of slots is two to the K. We're going to visualize it again. So, here is a postage batch.
On the horizontal axis, we have the 65,000 buckets, and since K is two, we have four slots in each. And now, here is some key information. Every chunk that you push to the network, or in other words, you upload to the network, they are going to use up one slot. Not just any slot, they are going to use up a slot from a very specific bucket. Okay.
So, now, we are going back to proximity order. Each chunk occupies one free slot, from the bucket, which is the closest to the chunk address. Or in other words, where the proximity order is 16. Okay. So, if we come back here for a second, now we can do a quick calculation that we have the 65,000 buckets, each has four slots, and each chunk is four kilobytes.
This is one gigabyte of data. But there's a very important difference. This is not actually one gigabyte of data. This is 16 kilobytes of data times 65,000. This is going to have a role just in a minute.
Okay. So, the previous one, this one was about the chunk and slot relation, and this one is about the node and the bucket relation. And here, I could replace the word node with neighborhood. In one neighborhood, nodes store the same buckets and same chunks. Okay.
So, we can go back to the rule with the store node and the neighborhood, and we can rephrase it in a way saying that every node stores chunks for some buckets where the proximity order is greater than or equals the network depth. This is the same network depth that we used for the network visualization. Okay. So, here we have depth four, meaning we have four neighborhoods. And the first neighborhood where the binary 00 neighborhood, where all the node addresses start with binary 00, meaning they store chunks where the address starts with binary 00, meaning those chunks are in the buckets which start with binary 00.
You get the idea. This might seem a bit complicated, maybe unnecessary complexity, but this has a very important purpose. That purpose is spam protection. So, let's come back here for a second and use our previous example that we have both one gigabyte of storage. So, let's say that I want to attack the network by overloading a neighborhood or a node with more data than it can handle.
So, I have one gigabyte of storage. The first neighborhood and any of these neighborhoods are only storing chunks for 16,000 buckets. Okay, one-fourth of my 65,000 buckets. That means that if I want to target this first neighborhood, I can only actually upload 250 megabytes there. In this wrong network, the network depth is at 10.
That means you would need to buy the thousand times the data that you want to upload to a single neighborhood to attack. The other purpose of the postage bag system and the on-chain payments is that it gets rewarded to the node operators, to the nodes which are storing the content. Okay, there is one last piece to the generating information part. That is the chunk types. There are going to be two of them.
The first chunk type is the content address. This is the simple one. Its address is derived from its payload. Okay, so if you change just one byte in the payload, you are going to get a completely different address. That means they are immutable.
If you fetch a chunk from the network based on its address, you get the payload. You hash the payload. Your node does this automatically, and the hash it receives or the address it receives must match the address that was requested. So, they are verifiable. These types of chunks are mostly used for plain old data that you want to put on the network.
The other type of chunks is the single owner chunk, and this has a different addressing scheme. It's made up of two parts. The first one is the owner or an Ethereum address, and the second one is an arbitrary identifier. This type of chunk can be used for conventions. I'm going to give an example.
Let's say I tell you to look up my homepage on Swarm. I'm going to give you two hints. The first one is that my Ethereum address is well known. You already know it. And the second one is that there is a convention on Swarm that people put their homepages at the identifier or zero bytes.
Okay? You can guess where my homepage would be on Swarm given these two informations. With content address chunks, you wouldn't be able to do it since you would need to know the payload. These types of chunks, the single owner chunks, are usually just pointers to content address chunks. If I upload my website, I do that with content address chunks, so my website is immutable.
And to create a pointer to it, I use a single owner chunk. One convention of single owner chunks is the feed. This is not a new chunk type. This is just a way to set the identifier in the single owner chunk, which used to be arbitrary. It is still somewhat arbitrary.
So the idea here is to add an index part to your identifier so you can version your content. We are going to see in the upcoming example that when you up your ENS with your content hash, ideally, you only want to do it once. But what happens if you realize that your website has a title that you fixed? You upload it again, you get a different hash, you need to set content hash again, so that's not so good. So the idea with feeds is to fake mutable data, to have a sequence of updates that you with your node can walk and get to the latest version.
So here's a quick visualization of it. You have the address of a feed, so you want to look up the latest version of the resource behind that feed. Okay, so you go to the first single owner chunk, which follows the conventions of a feed and see that it exists. Cool. That means you go to the next one that also exists.
You go to the next one, exists, and you go to the fourth index, which gets you a 404. So you go back to the previous one, mark that as the latest. So as we discussed, the single owner chunk has a pointer to the content address chunk, so you get that reference from its payload and go there and fetch that content address chunk. Okay, so this is how you can have mutable data where chunks are mostly immutable under the hood. Okay, that puts me at the end of this part.
Now I'm going to go back to my terminal, start the D-nodes. In the other screen, I'm going to monitor its status. I'm mostly waiting for the network depth value to stabilize. So if I want the initial value while the node is still being set up, it should be 10 after some seconds, which is the actual depth of the network. If you have any questions in the time being, just let me know and I will answer that.
I already have ENS here that I will be using, and we are going to use myens.cafe. If for now it currently has this dummy page. Okay, so I decided that my D-node is not playing nicely, so I'm giving it a restart. And now we are at depth 10.
Okay, so the first thing that I'm going to do is use a tool called swarm-cli, that I have it as an alias on S, and use identity create, give it a name, specify that I want a private key, not a v3 format. If I could type. Okay, so this is needed for a seed. This is going to be the owner or the Ethereum address part. I already have some postage batches.
They are a postage batch of stems. That's why I use the stem command. I'm going to zoom in a bit, and I'm going to do a feed upload, and then specify the path. Now it's asking me for a postage batch to use. Now it's asking me for the identity or the owner or the Ethereum address to use.
Okay, so two things happened. The first one is the immutable swarm hash that was created and uploaded to the network. So my website is accessible using this hash, but also a SOC or feed was created, which should be a unique and constant address where I can access my website. I'm going to try opening this, verify that it works. Nice, and I'm going to copy this address, go to ENF, edit my records, hopefully do the transaction.
I'm going to give it like half a minute, and it should go through. Okay, any questions in the meantime? Once we have ENF set up, we will try loading our website call through ethremal, and what will happen there is that ethremal goes to the swarm gateway, and swarm gateway goes to the hash, and it will be served even without a swarm node. Okay, so now I'm going to refresh this page. There's probably some cache going behind the scenes, so we're going to retry this just a few times, and then the magic should happen.
Unless there's like a five-minute cache. So in the time being, I'm going to go ahead with my demo, go back actually to the presentation that I've just shown you, and I will upload that to my feed, to the same feed. If I'm using feed upload again with the dist folder, I'm selecting a postage batch, I'm selecting the identity. Before I do that, I'm going to check on this. So we are writing the feed, and we should be having the same feed as before.
Address. If I refresh this, this is our demo. Okay, what's happening here? Still getting the same old blank page which was previously, so I'm going to check on this. This is up to date.
Okay, so maybe this will happen some minutes after the end of my presentation, but if you, ah, cool, finally. So if you go to thatcafe.limo, you should be able to access the material for this workshop, and you can also play with it and try the visualizations and simulations. Okay, thank you, and now the questions, please, if you have any. Without any costs, I mean, the network is incentivized, so the cost is spending the token to create the postage batch.
After that, what you get for that is that you can shut down your node and no participation is required. As for tooling, there are UI tools which are easier to use than CLIs and our installers where you can set up your B node through an UI, so you don't need a terminal for the whole thing. The frontend itself, so it has two parts. Since it's uploaded with the content address chunks, it is immutable, but we are using this concept called feeds, which kind of fakes mutable content. So if I uploaded a new version of the website, it would be accessible through the same hash.
Under the hood, it's a different content address, but the feed is the same. Single owner chunks are in a sequence, so you can walk them and get to the latest version. And if I had a new version, if I made any changes, I would not have to update my ENS, so that would also pick up on the changes. There are solutions to that. There is a format called Mentoray, which is a manifest, which is an arbitrary key value collection.
That's how websites are hosted under the hood. All the subparts are a key in the collection and the value is the corresponding swarm hash. You can go low level, even with the swarm CLI or with your own code, to only make changes in the manifest itself. Okay, so you are not changing any but one key in that to change one subpart. So you can do that to avoid re-uploading the whole thing, but chunks which have already been assigned a slot are not duplicated.
So you won't be paying twice as much if you did the whole re-upload. There are 8,000 nodes in the swarm network. How decentralized is swarm? I would say the only centralization point are the boot nodes that you need to first connect to for the gossip protocol and then discover the other peers. Well, not E3 more specifically, but you can try at ethswarm.
org, which is a front-end service for uploading content. So this works without the B node. It uses a proxied B node behind the scenes. So this is not decentralized, but the network behind it, of course, is. So you would create your postage batch here by bridging or swapping the token automatically and create a postage batch and then upload.
So this is one example for an app that you can use to deploy other decentralized apps. Are there any limitations to the contents of the feeds? Nothing I can think of. Please specify what limitations you have in mind, but nothing in general. East.
link, I guess many of these are under the same umbrella and work the same way. So, for example, there is also bzzlimo. They work by having a centralized node, which is used for the proxying and they bridge the web3 contents to the web2 so you don't need a node, local node, to access content from decentralized storage networks. What's the cost of running a node on Swarm? How long does it take to break even?
Okay, B nodes have multiple ways of running. It ranges from ultralight to a full staking node. For an ultralight node, you don't even need a blockchain connection. So it can just fetch chunks from the network. You can use that to download content, but it cannot upload anything since it doesn't have a postage batch, which requires a JSON RPC.
The system requirements for an ultralight node is minimal. Light node is the next, has a blockchain connection, but it doesn't participate in many of the chunks. I think you can do ultralight and light on an old Raspberry Pi, which runs there very well. If you want to participate in storing chunks, you're going to need 20 gigabytes of disk space because you would be storing 4 million chunks. And if you want to participate in the storage incentives, so if you want to get rewarded for storing chunks, you would need a good CPU because the storage incentive game requires you to hash all 4 million of your chunks and submit that on-chain to have a chance to be selected as the winner and get the token for exchange as the reward.
So if you want to play the storage incentive game, you would need a decent configuration. Of course, it still sounds very arbitrary, but you would be mostly interested in the CPU so it can hash all the chunks quick enough. How does Swarm compare to other storage? This specific use case, which use case, sorry, I would say, so Swarm is an incentivized network, so that's how it compares to IPFS, of course. And I would say, with regards to RVs, the use case there is permanent storage.
With Swarm, you can have multiple use cases. You can have temporary fives with just one week of TTL. You can go 10 years with some data if you want. So this depends on how you create your postage batch. Of course, the more tokens you spend, the more slots it will have in its buckets, but you can also tweak its TTL.
So you can have multiple use cases depending on the TTL. I would also add that Swarm has the bandwidth incentives built into the core system that I, it wasn't in the context of the talk, but the system is fully incentivized because when you have storage and you have the storage incentives, that's good, but you're going to want to push and pull chunks to and from the network, and for that, you will need bandwidth. And without incentivizing bandwidth, why would nodes ever give you back any data for your request? So that's also part of the network. Yes, it can serve videos.
Yes, so the single owner chunks are mutable. I would say feeds are mutable. At least that's what feeds take. You can fetch the single owner chunk. Well, okay, so you can do the hashing yourself with the owner, with the identifier, and then walk the indices like 0, 1, 2, 3, 4, and that would reveal the historic content addresses.
Out of the 8,000 nodes, how many of them are run by Swarm? Privacy reasons, not answering who owns what or who runs how many nodes. Do you have anything else?
Automatic transcript — names and jargon may be misspelled.