Securing Ethereum: Trust No Single Witness | Bhargava Shastry (Berlin Ethereum Day, June 2026)
Berlin Ethereum Meetup·Wed, Sep 9, 2026, 12:00 AM
Speaker
Ethereum has no ground truth but only independent clients that can disagree, so a bug is just a disagreement. This talk by Bhargava Shastry (Ethereum Foundation) covers how to manufacture those disagreements with nightly differential EVM fuzzing across multiple execution layer clients, how to triage the ones the world reports through the EF bug-bounty program, and why every oracle (fuzzer, AI, reporter) is fallible enough that a human still has to break the tie. The Berlin Ethereum Day was a one-day event held on June 15, 2026 during the Berlin Blockchain Week. The full-day program brought together speakers from the Ethereum Foundation and the broader FOSS, privacy, and security ecosystems to explore the future of Ethereum and self-sovereign technologies - from technical direction and core values to the challenges and opportunities ahead. Future Meetups and Events: https://www.meetup.com/berlin-ethereum-meetup/ More information on the speakers and the agenda: https://berlinethereumday.com/
Transcript
My name's Ivan, and I work for the protocol security and the Ethereum Foundation. Um I'm going to be talking about um the work that we do. Um essentially uh it ties two threads together. Internally, we we try to secure the system, uh by which I mean we do fuzzing, code audits, um look at uh denial-of-service attacks internally, but we also triage bug reports that comes to us from the external world through the EF bug bounty program. Um so this talk is about something that ties these two threads together, um which I'm going to be talking about.
Uh what makes us find bugs in Ethereum? Uh here's my talk in a nutshell. It is that there is no ground truth in Ethereum. I mean, we have a bunch of witnesses, um mostly on the execution side. You have uh execution layer clients like Geth, uh Nethermind, Besu, and so on and so forth.
And on the consensus side, you have clients like Lighthouse, uh Teku, Prism. All of these uh are independent implementations. There is no ground truth. Um I mean, we have uh partial ground truth, which is execute um executable specs both on the execution side and on the consensus side, uh but that's only partial. It's not a full-blown client implementation.
So it misses a lot of things like P2P, uh block validation, uh resource usage, and so on and so forth, which surface in real-world implementations. Uh so the core of the talk is how do you secure a system which does not have a ground truth? That gets to the definition of what a bug is. And in this setup, a bug is nothing but a disagreement. So once we have the working definition of a bug, um we can um essentially take it further.
So, uh uh, Ethereum bets on client diversity, by which I mean, uh, more the better. Uh, if one client has a nasty bug, it doesn't endanger the network as long as, uh, other client shares are, uh, big enough. Um, and they can keep the network running. But, it's the very property that keeps the network resilient, which makes security hard. Because, if two clients differ, um, slightly in how they implement consensus, then you have a hard fork, or you have a consensus split, and that is catastrophic.
So, essentially, we establish that the very thing, the very property that makes, uh, mainnet resilient, also makes it, uh, brittle in some ways. And, uh, so, when we say bug is a disagreement, uh, we start looking at techniques we can use to surface these dis- disagreements. And, one obvious way to surface disagreements is to essentially have them run against each other. So, you run, let's say, Geth, and say, uh, and notice what it produces, and then you run another independent client like Nethermind, notice what it produces, and then, uh, if there's a disagreement, uh, we flag that as a bug for manual triage, and then report it to client teams. So, this is one line of work that we do.
Um, so, essentially, the takeaway from this slide is, diversity is a blessing, but also a bane, because it makes our job, uh, a lot harder. We need to not only look at one client, if one client has been implemented correctly against the spec, but also if multiple clients speak the same language. So, let's talk about how we manufacture disagreements. One way we do that is fuzzing. That's a pretty big line of work in our team.
And, uh, we use at the execution layer something called Go EVM lab which was written by Martin Swender. Um essentially what it is is it's a fuzzing machine which essentially creates EVM executable code out of thin air. And it feeds it to different clients and then we look at the execution trace essentially a snapshot of what each client did as it was executing this piece of code. And we compare the final state root and hash of the execution trace. If they differ uh then there's a bug.
Most often it's boring. You don't see any any alerts going up. But occasionally, and I'm going to talk about it in the next slides, we we do see disagreements on production line not production line code but tip of master so to speak. Uh the clients that we test are geth, Nethermind, Besu, all all the ones that are big and also the minority share ones like with Nimbus uh EVM one and our own reference implementation called Eels. And the thing that makes this useful is that we run this on a nightly basis.
So every night we fetch the tip of the trunk of each of these code bases, build the EVM binary from them and manufacture test cases at a pretty high throughput. We have a machine which is quite beefy and helps us scale up. So we run multiple instances of each of these clients and we can get to a throughput of roughly 2 to 3,000 executions per second. So um compounded over 24 hours because we do this every day, gives us roughly an assurance of over 100 to 200 million test cases every day. So what I want to what I want you to take away from the slide is the fuzzer that we run is essentially CI for consensus.
It catches regressions uh as soon as the commit lands. Let me uh make that a little more concrete um by talking about a real divergence that surfaced last summer. And uh I can talk about it mostly because the disagreement never made it to production code, which essentially um talks about uh how important it is that we do this product uh proactive line of work that we do. Um so, this happened during Osaka fork fork prep uh last summer, around July. Um one of the clients, execution layer clients, uh ships a new uh BN254 precompiled back end.
And uh the bug was that this back end silently accepts curve points that are not on the uh points that are not on the curve and silently reduces them modulo n, which shouldn't happen, which didn't happen even with this client in the past, but that was a regression. Uh the other clients continued working as they were, so they would reject this point, and when a precompiled call was made with a point that was not on the curve, they would say, "Hey, this does not belong here." And uh failed the precompiled call, but the buggy client, because of the regression, because they introduced a new back end, would silently accept the point. And this was quickly flagged by the fuzzer because the fuzzer is essentially manufacturing uh contracts like these, so it doesn't take it too long. So, essentially, about 1:00 in the morning is when the buggy commit, the regression, was introduced.
About 12 hours later, we get a paging alert from the fuzzer. So, we write this. So, every time there's a disagreement, there's a ping on one of our phones, and then we can pass this on to the client teams. So, we get a a paging alert about 12 hours later after the the bug was introduced. We uh essentially notify the client team.
We have a coordination channel with each of the client teams. So, we use uh this to um share uh threat intel. And about 2 days later, the commit uh the regression is fixed, and this never sees the light of day. There's no production uh impact. So, uh the main takeaway here is that you have a client uh splitting disagreement here, which was found within 12 12 hours of the commit, and gone in about 2 days.
And uh the nice part is because we do this on a daily basis or nightly basis, uh the commit that was responsible to cause the dis- disagreement is usually staring right at you. So, it's pretty easy to triage and diagnose and bug fix, as against if it were found a couple of months later, then you need to go bisect the commits and uh do a more deep dive. And uh this is not a stroke of luck. The main reason this works is essentially reproducibility. It is that um we have a standardized EVM status interface.
So, there is an EIP that defines how um tests can be run. And there's also an EIP which um decides what trace each client needs to output when a test is being run. So, we make use of these standardized interfaces and essentially feed feed uh tests at scale. So, uh this pipeline, which is uh create a test, minimize it, uh um reduce it, and then uh pass it on to the team. This pipeline is what makes the whole uh exercise worthwhile.
And when we send a report to the client team, it's not a vague uh AI-generated report or something. It's a it's a file which they can run at their end, and that should produce the same output that uh we got, and the reason we got a paging alert, so it is reproducible and easy to work with. Um I'm not going to be talking about specific bugs it found, but uh the shapes of bugs it found, you have things like state root divergence because one client does something slightly differently or consume slightly more gas or less gas. And you have uh essentially with new EIPs that are implemented like 7702, set for transactions which came with the Prague hard fork. Uh you have edge cases which uh turn into bugs.
Now, here's the first uh turn of the talk, which is that so far so good. So, we we have uh a machine that can generate test cases and discover disagreements, but um only a subset of bugs is a disagreement. What if all clients that we are testing implement the spec the same wrong way? So, you have an agreement, but doesn't mean the implementation is correct because spec uh said something else. Uh this does happen uh although uh less often a lot less often than um clients uh differing in the implementation itself, but it's important to be aware of the blind spots of the machine that you run in order to essentially cover for the blind spot.
The other blind spot so to speak is networking attacks. So, so far I talked talked about how we fuzz the cons um the consensus core engine in the execution layer uh side, but then there are a whole bunch of attacks which uh use networking as a medium to either cause denial of service by amplifying resource usage in the client side or uh in more serious uh scenarios by crafting a block which would lead to a consensus split uh and or bring down a bunch of nodes during block validation. These have happened and these are known blind spots of the the fuzzing machine that I talked about. So, we are a team of close to 10 people. So, the fuzzing I spoke about in the last few slides is not the only thing that we do.
There are people looking at networking side attacks, resource applications, and so on and so forth. And we also run EI agents internally to look at new threat surfaces. So, the main thing I want you to take away is that the article that I spoke about, Go EVM lab, is blind to the most nasty bug, so to speak, which is all clients agree but agree wrongly. And also that we have deficiencies on the networking stack, and we have colleagues in my team who are looking at it. Exactly these blind spots.
So, one way we thought of covering for our blind spots is to essentially have articles that look at bugs from a different perspective. So, so far I've spoke spoken about how a bug is a disagreement, but what if we don't rely on implementation differences, but go back to the spec and derive properties derived from the spec itself. I'm talking about the the pros of the text, not the implementation, but the text of the of an EIP, and then extract properties, and then create test cases based on properties. So, these are the properties the spec says is valid or invalid. Do the clients actually agree with that assessment?
And if they don't, then you have a bug which is outside of the realm of client A versus client B disagreement, but spec versus implementation disagreement. Of course, we also have bugs from time to time which are very rarely exercised by the fuzzer. I mean, cryptographic code is uh a well-known blind spot for fuzzers. It's very hard to, for example, uh forge a signature. But, there are also other edge cases which make purely fuzzing uh problematic or deficient in providing a strong assurance.
So, for known blind spots, we have uh handwritten engines which cover for them. Uh we also um look at targeted fuzzers. So, if we receive a bug report through the bug bounty program, which shows that there's something that we don't cover, we try to write fuzzers for them. Um and try to cover the the weak spots. So, essentially, what I want to take you want you to take away is that the total assurance that we provide is internal plus external.
So, there's a bunch of work that we do internally at the in the protocol security team, but we also get threat intel externally through the bug bounty program. Let me talk about the bug bounty program uh now. That's the second part of the talk. Um so, proactive fuzzing that we do is internal. Uh of course, there's a bunch of other stuff that we do.
Uh then, there's this external line of work, which is uh the Ethereum Foundation bug bounty program. Essentially, the intake is uh it used to be a Google Form, uh but EI changed everything. Um so, I'm going to talk about that in the next slide. So, it's essentially some uh a form that an external reporter uh submits. Now, we require them to burn a certain amount of money, about $10, to uh submit a bug bounty report because we noticed that uh pre-EI and post-EI, the the channel has been flooded.
So, it simply was too hard for us to manually triage this uh quantum of reports. So, they do a small burn and send us a bug a bug report. And then, we try to essentially deduplicate it, make sure that it's not a known bug report from the past. And then, the most critical part is actually reproducing what the reporter claims is true. So, for example, if they say Lighthouse has a denial-of-service attack through this P2P networking vector, we need to actually make sure that what the reporter claims, for example, an 8-gig node coming down in 20 seconds is actually true before we can report it to the client team.
Um so, this is one of the most important things that we do as far as POCs are concerned that we get externally. And I like to make this point that we are broad. We're not specialists. The people who know the bug inside out and the fix are the client developers, but our job as generalists is to amp up the signal-to-noise ratio. So, we are the interface between external bounty researchers and the client teams.
If we simply pass on every report that we get, then it's going to be really noisy and it's going to lead to this problem that the client teams are simply over-burdened with a lot of bogus reports, and even the genuine ones get shadowed. So, we are the middle people, so to speak, who are doing the triage, severity assignment, and handing a small subset of high-quality reports to the client teams. Now, here's the second turn of the talk. What happened post-AI was quite essentially that we got a massive chunk of reports. So, we used to get maybe three to four reports a week through AI.
And then, we started getting like 60 to 70 reports every weekend. And a lot of it was AI-generated. So, people would spin up um, agents and simply report things that these agents found, which may or may not be true. A lot of it was AI slop um, claiming uh claiming that they would it could bring down the network, which on closer inspection turned out to be uh false positives. Now, uh imagine you're dealing with uh 10 to 20 reports every day, and it's simply not possible for a small team like ourselves to be uh dealing with this quantum of reports.
So, essentially, we used AI against itself. So, we also have AI swarms internally. And my colleagues have taken this a step further more recently, and uh we have a semiotic close to a semi-automated to full automated pipeline from the intake of a bug report to triage, severity assignment, and passing it on to the client team. So, um essentially, AI rates and reproduces the flood that AI creates. Uh but a human breaks the tie.
So, we sit in between. We don't simply pass on uh a report uh if it is approved by AI. We sit in between. Uh we know how clients work. We are generalists, so to speak.
We try to reproduce it. If we can reproduce it on the machine, then uh we pass it on. So, the human judge is still necessary. We are running AI against AI, but AI doesn't replace the judge, which is human sitting and making the final judgment. Now, I'm going to be talking about the signal uh signals that a bug report sends.
So, I'm going to be stacking up three different signals. One, what the reporter human reporter reports. Uh they could be using uh manual audit or using an AI and say, "I prove that uh get is going to crash by sending this packet." That is their claim severity. Then, uh Uh we have our own AI triage swarm, which says, "Okay, I'm going to run the POC that the reporter sent us and I'm going to use the rubric which is the market share of a client, how expensive this attack is, and what is the impact it has on the mainnet, and assign it uh a severity rating.
This I define as agent severity. And uh it can boil down to does it reproduce? Does it match the the the claim by the reporter? And then there's blast radius which is what is the impact of this bug on mainnet. So, um severity for us is it crashes.
I mean, yeah, it's easy to say that a packet can crash a client, but the end goal is to ascertain uh uh what is the impact of this on mainnet. If a crash happens, how bad is it? And how resilient is the network to a crash like this? So, it's um severity isn't isn't it crashes, it's amplification times propagation times asymmetry. Um for example, we had a recent report which claimed that uh there was a denial of service against a major client, but we found that the attacker needs to do as much work as the client does.
So, it's a one-is-to-one work, uh which quickly downgraded the report from uh medium to an info. Um so, essentially um honest downgrading of reports that we receive is also um uh a good skill to have in this line of work. We need to make sure that uh it's not just um a packet that crashes a specific node, but it it's also something that's impactful for mainnet. Now, here's the other thing about triage. It's it's a vantage point, not a cube.
What I mean by that is uh triage from uh of of reports from a bunch of researchers around the world gives us a very high-level view of bug shapes about uh what is going wrong today um at a high at a very high level. And we can use this info internally as well to improve our own assessment of clients, which we do on a uh on a timely basis. But, it also stacks up um the the difficulty against itself. So, client diversity makes securing a system difficult because we have 10 different client implementations. But, if we get a bug report that affects uh let say Lighthouse, we can use the same shape of the report and query if the same bug exists in Teku and Prysm.
So, this gives us um the same tools to look at different implementations. So, what makes client diversity tick can also help us uh in our favor when we look at bug reports. So, I'm going to synthesize the the talk that I've given so far. Essentially, it is that we do two lines of work in the security team. We do internal fuzzing, code audits, and so on and so forth.
And we do bug bounty triage through the EF bug bounty program. And one thing that ties these two threads together is that uh bugs are usually disagreements. Uh it could be between two client implementations, or it could be between uh the reporter, the AI that we run, which triages a bug report, and uh the the final ground truth, which is uh judged by a human. So, we can use AI to essentially scale scale this process up. So, if we know bug shapes, we can run agents to find more bugs of the shape.
But, we cannot replace the human in the loop because the ultimate decision uh how this bug impacts mainnet requires experience, requires uh broad uh expertise, which I feel at least state of the AI uh art AI does not have. So, um Yeah, three main takeaways are a bug is a disagreement, a green check mark doesn't mean that everything is okay. We need to look for blind spots. Uh so, build oracles that also help you catch blind spots. And finally, AI scales witnesses so you can generate more bug reports confidently, but it also requires a human in the loop to finally be the judge of what is problematic and what is not.
Thank you.
[applause]
Automatic transcript — names and jargon may be misspelled.