testing in production: gathering actually useful data
ETHBerlin·Mon, Jun 16, 2025, 03:08 PM · 17:48
I'll show a both novel and useful way to test blockchains both before and in production. i'll present how we test every both non and state-breaking change to critical components of our stack (consensus, state machine, communication protocol) and how we infer useful information from the data we gather before mainnets.
Transcript
I'm Jurgis, I'm the platform team lead at Intershare Labs. This talk was supposed to be about production, but then I thought it'd be more interesting to talk about how we do it before production. So, oh, the slides aren't there. Thanks. Cool, so I'll give some brief context about what context we work in.
So, I've worked in Cosmos for about four years now, and the constant wall, no matter what we work on that we keep running into, is how do we test any of our code? Cosmos, even though it uses the same underlying stacks, there's actually a lot of different both implementations, modules, and even forks in the ecosystem that all have various kind of nooks and crannies that you might miss when writing code for them. And then we're also a platform team of two, now three people. And I'm also the only, was the only DevOps engineer until a month ago, so this means whatever solution we reach for, it has to be mostly self-service, because we just don't have the capacity to handhold developers for whatever they need to do. And about a half a year ago, we were called Skip before, we got acquired, and now we're called Interchain Labs, which meant that all of these critical components that I mentioned before, now we're also responsible for maintaining them.
And the goal of the talk, I think similar to other talks about testing today, is not necessarily to show you on how we do testing or on our infra setup, but it's to show you how we experimented with the DevOps of infra, and to also just share our learnings about what works for us and what didn't, and kind of for you not to go the same path that we did necessarily. So before kind of the first wall of testing that we ran into was even like at the start of Skip, we had a block builder product that was kind of similar to Flashbots, where we had a centralized block builder, and then validators would connect to us for blocks. And what I mentioned before, where people had their own forks and stuff, we also had our own fork, specifically for the consensus engine, and because we supported more than one chain, whenever we did a release, this meant that we had to fork each chain's consensus engine, apply some patches to it, release it, and then validators would use those binaries. And in some networks, we had almost 100% coverage, which meant that if we pushed a bad chain, even though the fork contributors had nothing to do with it, we could crash the whole chain. So to test it, we would actually set up DevNets that had a few validators that were running our software, a few that were running the original software, to make sure that they're compatible, and the tool was called Builder.
It was just like the first thing I did when I joined Skip. It was a super hacky pipeline script that would generate some Docker Compose files. We ran them locally, and you'd see that everything's okay. And then a few years later, we inherited that kind of maintenance of the tech stack, and when we started looking into it, we weren't super happy with the existing testing flow. It had really good coverage of unit tests, integration tests, end-to-end tests, everything aside from actually running a network that represented at least something similar to a testnet or a mainnet, which meant that we weren't really sure that whatever we released out there for our customers, for our users, would actually work in production.
And then on top of that, we also became responsible for the security maintenance of the software, which is super scary, and then also, whenever we had to release a patch, users would usually upgrade really quickly, which meant they also didn't really do a lot of testing because there's just not so much time before a patch becomes public, and then you're kind of in a race with attackers. So we kind of sat down and thought about what we would want for our pre-production testing, and one of the main requirements was that it shouldn't replace any of the existing testing, so we weren't trying to make it more efficient or anything like that. It should be like net new capabilities for the developers. It should also be flexible enough to work for different use cases, so whether we want to load test a network or whether it is we want to test state compatibility, meaning that the network wouldn't fork if we released a patch, whatever we did was supposed to support all of those use cases. We also didn't want to simulate conditions, so I think in some of the talks before me, especially I think in the Polkadot one, they had a co-located data center where they would simulate latency if they wanted a condition similar to mainnet or testnet.
We didn't really want to bother with that because we also couldn't really know what those conditions were in the actual network, so we wanted something that represented a mainnet, but we didn't want to simulate it, so this meant maybe multi-region or big machines, whatever it meant, we just didn't want to simulate it. And then GoodWX is kind of a no-brainer for the developers to use it. We didn't want them to go write Terraform or Ansible. We wanted to meet them where they're at, and it also had to be fast, and I'll talk more about that later, but the gist is that we didn't want to wait 30 minutes or more to just see a network be live. And then we also took some inspiration from other people in the ecosystem, so eFandOps was also talked about today.
They use a mixture of Kurtosis scripts, which is a tool to set up local networks, and then a few other services for managing those networks. So this didn't work for us in terms of the WX point. We just did not want developers touching Kurtosis scripts because they're not super nice to work with if you don't have the knowledge of the DSL. Celestia, as far as I know, maybe not recently, but they used to use a fork of Deskground, which is also similar to Kurtosis, that sets up multiple services, and they can either be local, or I think they had support for the cloud in some ways, but it was just not flexible enough for us, and it was also, I think, configured, a lot of it was configured in YAML, which goes back to the WX point. And then there's another team in the ecosystem called StrangeBot, and they have a Kubernetes operator that just manages nodes in Kubernetes and does backups, snapshots, things like that, and we run Kubernetes, but we just did not want to use Kubernetes for running nodes.
And yeah, and the existing setup for testing was a bunch of Terraform and Ansible that would set up some amount of nodes on DigitalOcean. So we did not want to use Terraform for these two reasons, is that, one, it would take hours for the infrastructure to be brought up and then configured by Ansible, it was just a really slow process. And then the second point is developers just do not like Terraform. I kind of understand why, it's not a super nice language to work with, so we didn't want any of it to be in Terraform, one, because it's slow, and because it's bad UX. So the first kind of iteration of what we wanted to do was Petri, which was a homegrown infra framework, ready to go, that would abstract the concept of containers, DigitalOcean, droplets, whatever, into just two free concepts of a task node and chain.
So the chain would be composed of nodes, which would be run on tasks, and it worked quite nice for iterating on kind of one-time experiments. So whenever we needed to, for example, load test our on-chain Oracle or just run a network for state compatibility reasons, it worked, but configuring those networks using Golang and if there are kind of different variations of those networks was just a bit too cumbersome. And then we also didn't really handle the build step well, where the developers would also have to figure out how to build those nodes manually, which just also led to bad dev UX. So you can see on here, the setup looks simple, at least on paper, but there is just a ton of configuration parameters that you need to set up. And for each chain, they can be different.
And you need to be aware of all of these concepts as a developer. So we ultimately kind of moved away from it. And what we moved away into was this service called Ironbird. So in P3, we learned a lot of the developer experience was really, really important for whatever kind of testing we wanted to do. So the main goal was a Versailles-like experience.
And what that means, or at least meant at the time, was that we wanted to be able to spin up a blockchain that could be configurable to whatever the specific protocol team wanted on each PR. And what we did was we reused a lot of P3's code and we created a GitHub app that just on every PR, depending on the repo's configuration, would spin up a DevNet and it would provide the status, the endpoints. We also had monitoring out of the box. And then we also used Temporal, which is a workflow execution engine to kind of handle failures, because there was a lot of components in there like DigitalOcean, even our software that could be at times unreliable. So we wanted something to be able to handle those failures.
And this is what kind of the flow looked like. So whenever a PR would trigger, we'd create a workflow. We'd build the actual Docker image, which was a really nice step to abstract away from the developers. We created the droplets, the Docker containers on those droplets. Then we'd configure the chain, spin up a load balancer to expose the chain to the public.
So if a team wanted to, for example, demonstrate whatever changes they made to an external team, they could do that. And then the workflow after spin up would just continuously track the network health until it was shut down. And if there was anything wrong, we could either retry something or even just shut down the chain and try again, because in most cases that was something we could do. Yeah, so we also, for the people out there, we wanted it to be fast and we actually made it fast. We used pre-built images that already had everything except the actual node software on there, which meant we could get a DigitalOcean droplet up and running in about a minute.
And it also had that telemetry in the image so developers could immediately, whenever a node or a chain was spun up, could just go to our Grafana instance and see all of the metrics there and even compare metrics between runs. So, for example, if you were doing load testing, you could compare two Grafana dashboards or two different time ranges and see what the differences were. And then we also, for kind of debugging and convenience, we also exposed the Docker containers or the raw Docker engine through Tailscale, which is our internal network. So if developers wanted to change something manually, they could just exec into the container and make any changes they wanted. And this might be a bit hard to see, but this is what kind of the flow looked like.
We could just send a GitHub comment, do you start an ironbred network? And then a GitHub check was created with all of the information about that network. And as you can see, the UX is still a bit rough because as we kept building and building more features on top of it, so most notably in here is a load test. GitHub comments was just not enough of an interface to actually make any customizations to those networks. So we realized that we had to get away from the GitHub app flow.
And what we did is the primary contributor, Nadeem, for the service just wipe-coded the hell out of a frontend for the developers to interact with. And the flow is actually mapped quite nice into just one simple form that I'll show later, that the developers could either pre-fill the values or customize anything they wanted and the network would be launched. And then another thing that we learned was we had all of this infrastructure and because we didn't use Terraform, we had to keep the state ourselves. And initially we kept it in the temporal workflows, but because we didn't explicitly need just ephemeral DevNets, that also didn't really work out well for us. So we just moved that into a database.
And lastly, we started focusing on just supporting internal engineers. We just made the kind of thinking about the product itself much simpler than having to support all of the 100 different configurations that you might run into in the Cosmos ecosystem. Yeah, and this is what the flow looks like. It's not Brazil-like anymore, but it's definitely more flexible for the engineers. So usually what you would just do is just select a repository.
So because we have different protocol teams internally, all of them have a different need. So maybe the Cosmos SDK team wants to test the Cosmos SDK, whereas we or the Hub team want to test the Cosmos Hub. And you specify the shop, the version that you want, how many nodes and any modifications that you want, including a low test or a long running testnet. And we just handle the spin up for you. And this is what it looks like when it's been up is just a pretty simple view of the network where we have different full nodes in the network, we have different validators and then also below it, I cut it off, but we also have the external information about the network and we also have the monitoring links for it.
And this isn't kind of still the end, I think, for Ironbird, where we eventually wanted to move this, become kind of the backbone of integration work between our teams internally. So we also have a product called SkipGo, which does cross-chain bridging across in and out of the Cosmos ecosystem. So we also want automatically whenever a developer sets up an Ironbird network, we want to be able to configure relayers. We want to be able to configure the live clients, the contracts for them and any infra internally that we need to. So whenever we wanted to test some feature work either on the SDK or on the Hub, we automatically support the network internally.
We also want to be able to set up faucets, explorers. So right now we still have to do quite a lot of work if we want to demonstrate it publicly to other developers. We still need to do a lot of work, whether it's sending tokens or just like having Explorer, which we don't right now, which would be really, really helpful. And then lastly, an idea we've been experimenting with is working mainnet state and replaying transactions from live networks. So this is, again, the kind of simulation piece I talked earlier.
We don't want to be pushing synthetic load tests because they don't really tell us what we want about user patterns and user activity on live networks. So we want to be able to just copy whatever people submit on mainnet and copy that on our chain and see if it still handles that. Yeah, and then we learned quite a lot of lessons. I think, yeah, one is keeping track of infrastructure state is really hard. So this is what's kind of a gut punch because we specifically didn't use Terraform because it keeps track of the state and we still needed to keep the state.
And it turned out to be really hard. Supporting M plus M so that 100 configurations of chains is really difficult. It wasn't impossible, but was still something we decided not to do to simplify the service. And then, yeah, it's just fun to experiment with. And for UX, I think, especially in blockchains, the industry was for a really, really long time stuck with just Terraform and Ansible scripts for setting up whatever they needed to do, or at most, like maybe Kubernetes.
But I think we learned that it doesn't have to be that way. Just the important thing to keep in mind is whenever you do that kind of stuff, it's much better to start simple and then put complexity on top instead of just trying to be smart and automate everything that you can, because you might realize that you don't need to automate it or it's just maybe too difficult to automate it and you can do it manually. Yeah, thank you for listening. If you're ever interested to chat about infrastructure, you can reach out to my Telegram or email at zxentrachainlabs. Thank you.
Automatic transcript — names and jargon may be misspelled.