New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

25-Minute Solidity Fuzzer: Fuzzing Smarter, Not Harder

ETHBerlinThu, Jun 19, 2025, 12:35 PM · 1:14:11

Fuzzing and Formal Methods** are often seen as competing approaches to smart contract security. In this **hands-on workshop**, we combine insights from both, allowing participants to build a **minimal EVM/Solidity smart contract fuzzer in Python within 25 minutes**.

Transcript

Good morning, everyone. Before we get started, let me just get to know the audience a bit. Could you raise your hand if you are a developer? Okay. Pretty much everyone.

Who is using a security related role? A bunch of people. Okay. Who has read or written EVM bytecode before? A bunch of people.

Okay. So I think we will have like something for all of you today. This talk is about fuzzing. Let me first introduce myself. My name is Thomas.

I run a boutique security firm that is doing cross ecosystem fuzzing, formal verification business post deployment security. Today's talk is going to be very much about solidity and the EVM, but we will have general insights here that you can apply across ecosystems and that you can also apply to fuzzing other things than smart contracts. So why talk about fuzzing in the first place? One thing that I've noticed, I actually come from a formal methods background. Let's just get everybody on the same page.

So what is fuzzing? So fuzzing, like a general definition would be, we create random inputs, we feed them into some program, and then we want to see if they break stuff. That's a very basic thing. So why talk about fuzzing? So one thing here is over the past year, both formal verification and fuzzing has seen more and more adoption in Web 3.

But I've also noticed that there's a bit of inciting, you know, that there are companies that are building formal verification tools, and they say, oh, there's this, like, case where your fuzzer cannot find the bug. And there's maybe companies building fuzzing tools and they say, oh, here's this case that your formal verification tool will not find. I think we have to be a bit more grown up in security. Of course, like all of these tools, they are not magic. Each of them has shortcomings, and they don't really work out of the box all of the time.

So you have to know what you're doing. The second reason to talk about fuzzing is kind of more fuzzing specific, which is in theory fuzzing is a black box technique, right? You just take your smart contract, you take your smart contract, if there's enough inputs, if there's a bug, it should find it. However, in practice, this is not the case. Fuzzing, if you try to fuzz a program, a smart contract to actually break it, you have to understand what the fuzzer is doing, you have to understand what it's doing to the smart contract.

And kind of the main point of the talk, if you remember anything, is this slide is fuzzing in practice is not black box. I want to show you a lot of examples of why this is the case. So let's get started. So this is an 80-minute workshop. We will use most of the time actually building a fuzzer from scratch.

Why are we doing that? Because I believe the best way to understand something is to actually implement it. You can play around with it, you can see what works, what does not work, you can fix it, you can dive deep into some of the things that break. So we will do this. We will do it in Python.

Of course, like in practice, you don't want to implement your fuzzer in Python because it's not going to be very efficient. But for a workshop, it's kind of a nice language, so we will do it in Python. We will dive into some necessary details of the EVM that we need to build this fuzzer. But really, what I want to do is to discuss with you then what works, what doesn't, to fix it, to iterate on what we get. And by the way, I want this to be interactive, like if you have a question, just shout out.

I can't really see a lot of you because of the light, but just give me a shout, and we can discuss. All right. And that's mostly the last content-based slide I have. And I really want to get going on the code. So there's a GitHub repo up there.

There's also a QR code. If you have your laptop, you can follow along. But what I would actually suggest is if you watch me coding and you pay attention to my screen here, because then, like, we can just stay in sync and we can discuss stuff and we have, like, one source of kind of that we are working on. All right. So that's it.

Let's get started. So let me say one thing about this GitHub repo. So I've set up this GitHub repo, kind of building an EVM fuzzer in 28 commits. So after the workshop, you can go back here. You can just browse it.

You can look at the step-by-step iterations that it makes. For this workshop, because we have a lot of time, but not so much time, I'm going to jump in, like, at an intermediate step where we have some basic setup. And let me just go through those basics for you. So we are building this on top of PyEVM. It's a Python EVM implementation that's pretty well maintained.

And I've set up a utility function here that allows us to get a VM instance, which we just kind of set up a chain and then initialize the chain and return us the EVM virtual machine. Then we just delete the class for creating accounts that will hold the private key, public key, that allows us to derive the address also of the account. And then, so we will be fuzzing. I've kind of created a sample contract here. This is a Foundry project, so there's a Berk token contract here, which is kind of an ERC20-ish token contract.

Since we're fuzzing, we don't really want to go too much into the details of what's implemented here. So we just have this utility function that will load the contract from the Solidity compiler output, deploy it on the chain. We have an assert here that asserts that the code at the contract address is actually the bytecode that we loaded from our Solidity compiler output. And then just as a sanity check, I have like two invocations here that invoke two basic functions on the contract, minting to an account 1,000 tokens. And then we query the balances and assert that we are actually just printing the balance so we can check whether the minting was successful.

So let me just bring up a terminal. And first thing we want to do is actually switch into the contract directory, just make sure that this contract is built. And then we can run this main file. And by the way, we will basically stay in this file and just see what's happening. So we're printing the account here.

The account has some address. It starts with a serial launch. Then we use this account to deploy the contract. This is the contract address. And then we are doing those invocations on the contract.

So after deploying, we are at nonce one, right? Then we do mint, which takes us to nonce two. Then we are querying the balance, invoking the view function. And indeed, the balance that we fetch from the smart contract state is 1,000. All right.

So far, so good. So we kind of have the basics here for fuzzing this contract now, right? We can deploy the contract. We can invoke functions on it. So let's start building our fuzzer.

And the first thing you want to do with a fuzzer is you want to check which functions you can fuzz, right? So let's create a Python array here. And we will actually, so when we are loading, when we do this get contract thing, what it will do is it will fetch from the Solidity compiler output the ABI of the contract. So the ABI is kind of the interface description of the contract. The Solidity compiler gives us this in as a JSON file.

It just says, okay, there's a kind of arguments. There's a function that has these kinds of arguments and returns these kinds of types. And this is like the basic building block that we need for building a fuzzer. So here we will just build this Python list of all of the functions that the contract exposes. And then what we want to do in the fuzzer, like there will be a big loop where we are kind of doing rollouts.

And let's say we are doing rollouts up to 1,000. So we will do like 1,000 invocations against the contract where we first randomly select a function call, right? So this is not, Copilot is trying to play games with me. But basically what we want is, yeah. So we randomly choose a function from our list of functions, right?

And then we are going to invoke this function. I think what's missing here is an import. Random. Good. So we have picked a function.

Now we need to generate inputs. That's kind of the other main thing. That's the main thing that the fuzzer should be doing, right? We have to generate random inputs. And since all of this is, EVM is working at the byte level, we can just generate random bytes.

Invoke the function with those random bytes and see what happens. So let's do that. Let's guess a number of the length of the byte array that we want to use as input. And let's be generous. Let's say we generate arguments between zero and 96 bytes long.

And then we generate arguments, right? So Python has this wonderful random library that allows us to generate a sequence of random bytes of this length. And now we can just invoke the contract. By the way, this computation object that's returned here by PyEVM allows us to check for errors, whether the contract reverted. It will give us this kind of stuff.

So it's useful to carry around. And one thing that anybody who has seen EVM contracts, like at the bytecode level, knows that the way we ‑‑ so each EVM contract has only a single entry point. You cannot really call a function directly. There's a single entry point. And what the solidity compiler does is it dispatches a table at the beginning of the contract.

And the dispatch table kind of dispatches based on a catcher cache of the function signature. Very low level. Maybe kind of weird, but that's the way it works. So we have to compute this function signature. And the function signature is basically ‑‑ let's say we do a mint, right?

So the function signature would just be the function name. And then it takes the types of the arguments. This is kind of the signature in text form. So we are doing this here, right? We're just constructing the function name, open parentheses, then we are comma joining the types of the arguments.

Then we take a catcher cache. And then the dispatch table or this signature fragment is just the first four bytes of this catcher cache. So we have technical implementation decisions, but that's how ABI compliant contracts work. So that's what we have to do. All right.

So we have a signature. And the arguments and the way we invoke this is ‑‑ so we pass our VM instance here. We say who is signing the message. And we will just use the same account that we used to deploy the contract. And we send this to the contract's address.

And then we have to specify which data we append. And the data here is just going to be the function signature concatenated with the random arguments that we created. Okay. And that's ‑‑ now we can like remove all of this sanity stuff here. And let's just do ‑‑ for debug reasons, let's make a print statement here so we know what our is doing.

This is not very helpful. Let's print the episode. And then let's print function name. And then let's print success. If ‑‑ oh, this is weird.

Success. Let's print success if the call was actually successful or failed if our function call reverted. Okay. So far, so good. We have built a fuzzer.

Like I promised, 25 minutes. Let's run it. Does it work, though? Right? That's the big question.

Let's invoke it. And let's look at what is happening. So what do we see here? Like we have a lot of failed calls, actually. Right?

It's weird. You can go looking. Like basically the only thing ‑‑ the only successes that we see is on few functions that don't take any arguments. This is like if you would dig into the contract source code, into the solidity source code, this is what you would see. So like all of the function calls here failed, which is not good.

Right? And this is kind of like one of the main obstacles that you encounter when you're fuzzing. I've seen fuzzing campaigns, even by security firms where people claim like, oh, we ran it for 5 billion runs. But if out of those 5 billion runs, all of them ‑‑ all of the invocations revert, they have not actually exercised the contract. They have not done a lot of work to find bugs.

So what can we do here? Let's figure out what's going wrong. And for this, let's ‑‑ in addition to the function name, let's print the arguments. I will just put this before here, because it aligns a bit nicer with the function name. And then let's print arg.

hex. That should give us a nice hex printout of our arguments. Let's run this again. So generating something looks like we intended, no? Like we said, the fuzzer is generating random inputs, throwing them at the contract.

But we want to see some successful invocations, right? We want to see successful transfers. We want to see successful minting of the token. Something that actually does something. Otherwise, our contract remains in the state right after deployment.

So what we can do here, let's grab a specific function that does not work. And this is let's look at balances. So we keep invoking balances here with kind of random bytes. I'm just interrupting this so we make progress. But you can run this for five hours and you won't see a successful invocation.

Why is that the case? Let's look at the balances function in the contract. So this is the balances function. But what the Solidity compiler does is for this mapping, it will create a view function, right? That takes an address as an argument and returns the associated UN256.

So this function takes an address. Let's look at the arguments that we are generating again. So we are kind of like generating bytes here, generating bytes here, generating bytes here. None of those are a valid address. If you know how ABI encoding works, an address in EVM is actually 32 bytes long.

And it just happens that none of the things that we pass here is 32 bytes long. So we could fix this, right? We could say, okay, let's not sample the length of the arguments. Let's just focus on this one function. We just want to pass this one function.

And let's fix the size of our arguments. It should work, right? I told you. Expect 32 bytes. Let's run it again.

Lots of failures. Anybody has an idea what's still going on? Yes. Yes. So an address in Ethereum is actually 20 bytes.

And when we submit it as an ABI encoded argument, the leftmost 12 bytes will be zero bytes. And if you look at this, like the way we are doing it, we are sampling the bytes uniformly at random, right? So we get all of this garbage on the left where there should be zeros. And it's very, very, very unlikely that we could use 12 consecutive bytes, which means like 24 consecutive zeros here, just because we are sampling uniformly at random. So it will give us some randomization in this way.

I hope I get this right. Byte strings in Python are always like a bit weird, but I think this should work. Let's pretend 12 zero bytes and then do the rest randomly. Let's run this. Ah, now it actually works.

But you see just generating random inputs doesn't work. There's some, our EVM contract requires some kind of well-formedness of the inputs. And without producing these well-formed inputs, we have no chance of invoking anything on the contract. And this is actually, if you think a bit about like go into fuzzing theory, we just see a major shift in fuzzing paradigms. So there's something called stunt fuzzing, which is what we've done before.

You just generate byte arrays, throw it at the program, see if it crashes. This is what works for a lot of programs and what worked very well in early fuzzing, right? When people are trying to crash image viewers. A PNG image is a bunch of bytes. But if you're trying to fuzz something that has kind of input validation, it doesn't work.

And smart contracts are one instance of this. So what we have just done is we switched from dump fuzzing to something that's called structural bare fuzzing, where we take into account what kind of shape our inputs have. Right. So with this insight, we can go ahead and we've kind of focused just on this balance function for now. But we can actually go ahead and try to fuzz all of the other functions as well, right?

They will need similar machinery where we are generating valid inputs for them. So to make this a bit quicker, so we'll just define this helper function here. I will just define this helper function here. Random input that takes a, let's say ABI type that is a string and returns the bytes that we generate for this. Actually, this should be.

Yeah, let's say any. So we return Python types with ABI and code them to get bytes. So if ABI type, and we'll abbreviate this, we will not cover all of the possible ABI types because that is a bit too long for this workshop. But what we will do is we will just like implement some of them. So what we can do, so we have this account type, right?

And what it will do is if you don't pass any private key, it will just sample 32 random bytes, take the catch up, maybe not super secure, but like fuzzing, for fuzzing it's fine. It will generate the private key for us. So to get a valid address, we can actually just create an account that will create a random private key. Derive the public key from that, and also the address. And then we need some other types that occur in the contract.

One of them is uint8, which is, this is very nice, of copilot, yes. The other one is uint256, so we just generate a random uint256. And the last thing is a byte 32. And for this, we can actually do what we did before. We generate 32 random bytes.

And if any other type comes along that we are not handling, we are just using an exception. And now down here, where we were generating the inputs, we can actually say, so we can get the input types from the contract's ABI. We can then, for each input type, we just generate, we call our utility function here to generate a random input. And then we can, from this, actually encode the call data that we will pass to the contract. So let me just check.

We have ABI. This is called encode. That's very good. And then we're computing the signature. So here we should be passing signature plus call data, right?

So now we are like, for each function that we encounter, we should, we should be generating a reasonable input. Okay. And let's just print, not this copilot, with inputs, random inputs here, why not. Because this, let's just, so we don't get byte garbage. Let's just say, use instance of bytes.

So we just hex format this thing. Oh, I'm sorry. Okay. Very nice. Let's run it again.

Maybe not grep for balances, but let's just pipe it into less. So we see we have some success, right? We are calling mint. We're calling staking shares, balances, all of these return with a successful invocation. At the beginning, we have a failed transfer.

That's not surprising, because we have not minted any tokens yet. This permit is failing. That's not very much about permits, because doing signatures in a fuzzer is kind of difficult. If we have enough time, we will get there. Transfer from is connected to permit.

Transfer from actually assumes that we did a valid permit before. So far, so good. Stake doesn't work, but this account might not have sufficient funds to stake. So it kind of looks reasonable. What you will notice, however, at some point, let's look at transfers.

So this first transfer fails, this one fails, this one fails, this one fails. And as we go on, right, we would expect at some point that our fuzzer mints some tokens to someone and somebody can do a transfer. Actually, let me say one thing before. We're doing one thing here. We're generating, whenever we see an address input type, we're generating a random address, kind of.

This is ignoring some things. It's ignoring the zero address, which you might also want to fuzz. It's also ignoring the addresses of precompiles, which, if you've paid attention, there was an exploit a couple of months ago where somebody circumvented the signature verification by passing the address of the identity precompile. So all of these things are kind of assumptions that we make when we are exercising the contract, when we are building or using a fuzzer, and we have to pay attention to those, because we might be missing things. So we might put it to do here, say, also fuzz 0x00 and the precompiles.

Okay. So why, as I said, we are not seeing successful transfers. Why is that the case? Because we are kind of like generating random accounts here, right? We might start a call to mint, create a random address.

We are minting to this address. And then the hope is also still like we are sending transactions still from just one account, right? So we would have to magically hit the address of that account, mint to that account, and then do a transfer from that account. It's kind of tricky. There's so many addresses.

You know how many addresses exist in Ethereum. It's very improbable that we by chance hit the same account again. So let's fix that as well. And the way we can do this is we can just create a small list of externally owned accounts. For all of these, we will generate the fresh address, write the fresh key pair.

Let's say we consider five. So we will consider kind of a small universe of accounts, and we will pretend that those are all of the accounts that matter for our fuzzing campaign. And we pick one of them as the deployer, as the address that employs our contract. Let's just say it's going to be the first one. And then we do the set of all addresses.

Why do we need the set of all addresses? Why would we? We will do something with it later. Because currently, so we will, let me say this, when we are sampling random inputs, we should return here one address from our addresses. But this list is ignoring one very important address that we also want to pass sometimes in our first course, which is the address of the contract that we deployed.

So we actually have to do, when we deploy the contract, we should append to this list the address of our contract. Otherwise, we are never testing the scenario where one of the inputs is the contract address. Okay. Then one thing we can do here, we can get rid of this argument. We now have a global variable of who is the deployer.

And then down here, we can also get rid of this. And instead of, actually, let's remove this. This we can remove. And then instead of this account, we just use our global deployer. And then for each function that we fuzz, we could use the deployer.

But actually, we want to also randomly sample who is invoking the contract. So let's choose from all the externally owned accounts. And let's use this to invoke our contract. So now we have five accounts. We will see them sometimes as input arguments.

We will sometimes see them as the sender of the transactions. That's kind of what we wanted. So let's look at, whoopsie, there is, that's what happens when you do live coding. 124, deploy contract. Wow, I'm missing right here.

So this is the deployer address. And also, okay. Now, by the way, like when you deploy a contract on Ethereum, like this, the address of the contract is deterministic at the point when you deploy it based on the norms of the accounts that does the deployment. So we can pre-compute the contract address here. That's what this line is doing.

Just a technical detail. So another fuzz run. Let's look at those transfer calls again. So we have a failed transfer. You can see we are passing an address here.

We're passing some amount. We have another failed transfer. We have another failed transfer, transfer, transfer. They're kind of all failing. And if I go, I don't know how far I can go here.

You will see that they are all failing. It's kind of weird. It's like we are now passing correct inputs. Why is it failing? And this is, again, the point where you have to go and actually look at what the fuzzer is doing.

You look at what's happening inside the contract. But we will fast forward here and I will kind of give you the answer. And what's happening is that we are not minting any tokens. Sounds curious. Why are we not minting tokens?

There is, this is the point where we should maybe look at the contract a bit. So as I said, this is like kind of a vault contract we have. It kind of contains a token. We are keeping balances of the circulating tokens. We have admins that can do admin stuff.

We're keeping track of the total supply of tokens. As I said, we are also doing permit style transfers. That's not too important. But we also have some vault-like functionality that we also don't have to care about for now. Let's look at mint.

So mint looks very unsuspicious, right? We are passing an address and an amount. And if it's not a zero address and it's not a zero amount, we are creating tokens. What's going on? So this mint function is of course only admins in our protocol should be allowed to mint tokens, right?

We have this only admin modifier. We require that the sender is an admin. And then if you look at a lot of function calls, if you look at a lot of rollouts, you will notice that there's this remove admin function. By the way, at the beginning, like we create one admin. The user that's deploying the contract is an admin in the beginning.

But what will happen, because we are sampling from the functions in the contract, there's a very good chance that we sample at some point this remove admin function. And the only admin that we have in our contract state loses the admin status. And all of the calls to mint will fail in this modifier. So there's kind of a race condition between mint and remove admin. And you might see one mint, but since they are equally likely to be sampled, we will see a lot of runs where remove admin happens first.

and we see no mints, or we see a single mint, and then remove admin happens, and then, okay, maybe we see some success for transfers. So another thing we have to fix kind of, we have to work around, we want to see transfers and it doesn't help if most of our fuzzing runs don't generate any interesting behavior. So what we might do in here is we might fix it by saying if function name is not constructed, but it's remove admin, and we just continue, we don't call remove admin ever. Let's look at our transfers again. So we will see some failed transfers in the beginning, because, of course, like we have to invoke a mint first, but at some point, if it's working, we should see a successful transfer.

Let's see if there's a live coding effect again. What I'm just going to do right here is let's work from a non-good commit. So I will jump into the point where we, okay, skip fuzzing, so we can jump in here, let's invoke previous fuzz. Ah, yeah, there's one thing that I did that I have not yet done with Q here. Let's remove this.

Okay. Previous fuzz. And... Successful transfer calls. There is a lot of stuff.

So we've seen this, like, the zero prefix with the addresses, that was kind of weird. We've seen now with this remove admin functionality, there's a lot of weird stuff that can happen when you do fuzzing that deviates from what you would see in normal contract behavior, right? Like when users normally interact with your contract. And one thing I want to raise is, is this way in which I included remove admin here, is this a good way to fix it? And it sometimes helps to not stare at the code, but to actually plot something, right?

So what we can do is, like, we can just count how many admins we see in our runs. And what will happen with this kind of code, so let's go back to the slides for a second. Right. So what will happen? So I've plotted here 10 runs of the fuzzer, right?

The x-axis is the depth of our exploration, right? So it's basically time. And on the y-axis, we see the number of admins. So what is our fuzzer doing now? There's an add admin function.

So it keeps adding admins. It will saturate at some point, because we have fixed the set of accounts to number five. It will never go down. The only thing we will see is, like, more admins. There's no admin removal.

It's kind of weird, because we have all of this unexplored space, right? We have all of this lower right part is unexplored. Maybe there are bugs right when there are no admin users, right? And not all of the users are admins. You could fix this again.

This is where fuzzing gets really, really tricky. You could fix this again. I was like, okay, this is maybe not realistic, right? And by the way, so the tricky thing in fuzzing is that you want to cover a lot of runs that exhibit normal contract behavior. But you also want to see runs that kind of exhibit exceptional behavior, right?

So you want most of your runs to have, like, some number of admins, not all of them, not none of them. But you want to see some runs that also have scenarios where all of the users, all of the accounts are admins, and you want to see runs where none of the users are admins, because exploits might also lurk there, right? You have to kind of explore all of it. Even if on those runs maybe other things are not as interesting, right? You might not do transfers.

You might not do minting, whatever. So how could we fix this? I will just show you a diff, right? So I went in. I was like, okay, let's just define a helper function that counts the number of admins.

And instead of skipping remove admin all of the time, let's say if only one admin is left, we keep that one admin. We don't invoke remove admin. If we fix the issue, we can go back, plot things. Plotting things is great. So now it looks like this.

It's kind of like more realistic, but we are still not covering this one scenario where we have no admins, right? Basically our fuzzer is very fast to promote some other user as the second admin. And then we said we never go below one admin. So we might still be missing bugs in this area where you don't see any colorful lines. There are other things that can go wrong in choosing basically distributions that we use to sample our inputs.

There's a fun thing. So I went and wrote this fuzzer, right? And at some point I was like, actually, I could show it to you, but let me just tell you because it's faster. You might look at it. So let's see if we ‑‑ here we actually have a number of mint cars.

Let me just output this into a file. Open the file. And let's look at the mint. The thing is we have to run it for some time. So you will see the effect.

But maybe I just tell you and then we can check back and you can convince yourself that this is the case. So if you run this, and by the way, if you were to write a real fuzzer, of course you would like encapsulate this entire main function into another loop, right? But at some point we reset, we start with a fresh VM and we start from scratch. We're just not doing this here because it becomes very convoluted and it becomes hard to show stuff when we have these two nested loops. But so one thing that can happen here is, and you will observe this on a lot of runs, is that you have one successful mint car and then all of the successive mint cars fail.

But the other thing that you might see is what we are slowly seeing here. You see successful mint cars, but the amount that we are minting is getting smaller and smaller and smaller. You will see that these amounts are decreasing in order of magnitude. So what's going on here? The thing is, again, we are sampling, it's again about sampling from our input distribution.

If we look at, I was really confused the first time I saw this. I was like there's a single mint and then there's no other mint happening, right? I fixed the admins, like what's going on? The thing is, so if we are sampling this UN256, right? And we will see an input distribution kind of like the diagram that you see on the upper right, just uniform distribution.

The thing is that half of the mints, right? So like this portion, the right half of the diagram, 50% of all of the values that we generate are bigger than half of the largest UN256 that we can represent. So when we sample, let's say we sample the first amounts for the first mint, and it comes from this upper half. And there's a high chance for it, right? There's 50% chance that we sample from this half.

And then we try to do another mint, and we sample an input again, and the amount is again from this half. What is going to happen is that we have two very big numbers. And when we compute the total supply, when we update total supply and add the second amount, it will just overflow. Because any two numbers from this upper half of our distribution, if added, will overflow during 256. In other words, like if you look at this lower diagram, I'm plotting kind of the log two, right?

So the bit length of the sampled integers. We have a lot of, we have this like long bar at the very end, so we're kind of like sampling very many, very long numbers. And we're kind of undersampling small numbers. And there's again, like you might want to fix this, or you might at least want to branch somewhere in your fuzzing campaign. You'll see like, okay, I'm fine.

I'm fine minting very close to 2 to the 256. But you might also want to branch where you say, okay, I'm only generating numbers from zero to the 100 or something. So your total supply, you see more realistic mints kind of, right? You see smaller mints that go to various accounts. Okay.

Good. So with that in mind, and let's just say like, we take it as it is. We accept that there's a big mint at some point, and then the mints are decreasing in size. What we really want to do is find bugs. And to find bugs in a fuzzer, you need to write some kind of property.

And so what we will do here is like, it's always a good idea. No. It's always a good idea to first. It's always a good idea to first state something that you know to be wrong, right? You expect it to fail just to like check whether your fuzzer is actually doing the right thing.

So we might write this utility function here, total supply, that invokes a Q function on our contract. Total supply returns the total supply of tokens that are in circulation. And then we just say it equals zero, right? And we expect this to fail because at some point we expect to see a mint, and this invariantly fails. So this is kind of just a sanity check to see that our fuzzer is actually operating.

So we can run this. And almost instantly you see, you know, we have a mint, total supply increases, and our fuzzer says, like, this property is violated. Amazing. So now the obvious thing is to write an invariant that does not hold. And for this, coming up with good invariants is always a tricky thing.

For this, we actually want to look at this contract a bit. So let's open the BERC token contract. So as I said, kind of an ERC20-ish contract. It's not implementing the full ERC20 interface, but we can mint tokens, we can transfer tokens. We have permit-based transfers, which we will not go into right now.

But there's a permit function and a transfer from function that allows us to basically define allowances by submitting some pre-generated signature. And then it has a vault-like function where the user can stake and unstake tokens. And as it is in Web3, right, we're always rushing to deploy, to release, right? There's, like, security is a second thought. So let's assume the stake function here is not completely finished in the sense that there are no rewards.

This is just to make this contract, to keep it at a size where I can use it as an example. So rewards have not been implemented, and whoever wrote this contract made kind of a simplifying assumption. They implemented the stake and unstake function, but they made one assumption, maybe assuming that the contract would be upgraded at a later time, right? Rewards would be implemented at a later time. But if you're skilled in kind of smart contract auditing and this kind of stuff, you might get suspicious by this line, right?

Because we have a division here. We are doing it, the division in the precision of the token. So we might have some precision loss here. The reason why we assume that this is not the case is because we are minting one share per token. So for any token that is staked, we will mint one share, and when a share is redeemed, we will return that number of tokens.

So as I said, there are no rewards implemented yet. But we are kind of making an assumption here that the number of shares that we are minting equals the number of tokens that are staked. And so we assume that this ratio, right, the price of the shares, is always one. But these two amounts are the same. And we can check this.

We can just write an invariant in our fuzzer to check whether this holds. So we will need another helper function here. Okay. And in this helper function, we will just query. So the invariant that we want to state, right, this is your basic invariant that we know that was failing.

So now let's write another invariant. Actually, let's first write the basic conservation law where we say that the number of minted shares equals the total staking shares, right? That's another thing that we assume, that the staking shares summed up over all of the users equals the total staking shares. So we can write just two helper functions to get, to invoke the few functions that will query this state for us. Let's see what Copilot did here, whether this makes sense.

This actually looks reasonable. Okay. And total staking shares doesn't take any arguments. Also looks reasonable. Okay.

So let's state our conservation law here. It will say that when we are summing up over all of the staking shares of, let's actually take, right, we want to say this for all of the externally owned accounts. The number of staking, the sum of the number of staking shares for each account equals the total staking shares. And we have to fix this printout just to make sense. Staking shares summed up, right.

And then we quit, if this is ever the case. So let's run this one. And the nice thing I told you before that we can, we could, like we would in a real fuzzer, we would have an extra loop around the kind of scroll out loop that we are doing here. What we can do instead is we can use, just use GNU parallel to run 10 of these Python scripts in parallel. And whenever one exits with exit code one, it will tell us, it will stop.

And we just output the output of all of these into a log file. So we can run this for some time. And I guess we should be pretty confident that our conservation law works. There's not too much going on here. We are incrementing the staking shares of the user and we are incrementing the total staking shares.

On the negative side, okay, we have some unchecked arithmetic. So we might actually see underflows here. That's a good thing to check this, right? Because who knows what's going on here. But we can keep running this and our invariant will not fail.

So now let's actually go ahead and add the invariant that I was talking about before, which is checking that the price of our shares is always one. So to do this, we actually already have, we'll add one more helper function that invokes a few function here, which is balance off. This is not right. This should be balances. As I said, we are not implementing the full ERC20 interface.

We're just clearing the few function of the mapping here. This is balance off. And then let's write a second invariant here. Where what we are saying in the invariant is that the balance of the staked tokens, so the way staking works in our contract is that we take them from the account that is invoking the contract and add them to the contract balance, right? So we are making a transfer from whoever is invoking the contract, transfer that balance into the contract and let it stay locked until somebody unstakes them again.

So we will query here the balance of the contract. And then we state that this should be equal to the total staking shares. So the total number of shares that have been minted. And if this invariant doesn't hold, we print an error message that this is not the case and we exit. And now let's run this thing again.

Whoopsie. So something failed here. Let's look at this log. What's going on? I'm only printing the successful invocations here because that's the only thing we care about.

Actually, when you're fuzzing, at some point you will exclude all of the pure functions, all of the few functions, because they are not changing contract states. They cannot drive your contract into a bad state. So what are we doing here? We're listing two mints, listing some add admins. And then the card that violates our invariant is a mint card.

How can mint violate our invariant? Yes. So let's check. So the arguments for our mint card here, this 0xE5 something, is actually the address of the token contract. So we are kind of giving, we're kind of like, through our mint card, we are staking without calling the staking function.

Now you might say, okay, mint is an admin only function, right? It's basically, this is the admin shooting themselves in the foot. If you do an audit competition, people will say, ah, it's user error. It's not a real finding. But let's prevent our admin from shooting themselves in the foot.

Let's add a require here, where we say that two must not be the contract itself, right? Because otherwise we are inflating. If the admin mints tokens to the contract, we are kind of like inflating the share price. And let's quickly change into the contract directory, run forge build, change back. Now we believe we fixed it, right?

So we just run our fuzzer again to check whether we actually fixed it. Everything is good. This should run forever, not find anything. So the fun thing about fuzzing is that you can basically parallelize it indefinitely, right? If you have as many courses as you have, you could run it in the cloud, and it just speeds up things.

But there was something again. Again, something violated our second invariant, I guess. What happened? Yes, it's again our second invariant. And this time it's a transfer.

So what is happening here is similar to the first thing. So we have a transfer here. And if we check this address, 0x98, it's again the contract address. So now we have an account that held tokens and transferred tokens to the contract. It's a donation-based attack.

So now it's no longer an admin error because anybody can transfer, right? And what this allows you to do, if you check in the repository afterwards, because I've also included POC in the foundry directory. So what this allows you to do is an attacker could transfer tokens to the contract, donate tokens to the contract, inflating the share price, and by staking a very small amount, exploit another user that stakes a much bigger amount. So what can we do? We can, again, fix this in our contract by also adding a require here.

This does not make any sense. And again, change into the contract directory, forge build, change back, run our fuzzer. And now it will run a very long time not finding anything. Are we done? No.

But I will, again, fast forward here for a second because we only have 10-ish minutes left. So I will just show you what we are not finding yet with the fuzzer. And what we are not finding yet is actually to do with these permit-based transfers. So basically what we do here, right, we allow someone, the spender, to sign a message that allows somebody else to invoke this contract, passing the signature, and spend money on behalf of the spender account. The issues are not here.

The issue is actually that in this transfer from function, we have the same issue that we had above with transfer, where we are not preventing transfers but donations into the contract. The reason why our fuzzer is not detecting it is that here we are checking a signature. And basically what the fuzzer would have to do to generate a valid signature here is invert the cryptographic function, right? It would have to guess the right input to pass as a signature so that it actually is a valid signature for this digest message. And we all believe that this is impossible, right, in finite time, in the time that our universe lasts, until quantum computing comes along.

We all believe this is not going to happen. So it's actually impossible, if you believe that, for a fuzzer to do this. And the only way to fix this is you can check in the GitHub repository. The only way to get past signature checks with a fuzzer is to actually pre-compute the signature. There's no way.

You cannot just generate random signatures and hope that they are somehow magically signing a message that you try to pass. So you actually have to kind of implement a lot of codes. I'm doing this here, right? I just have this PermitSigner class. Let me make this a bit bigger.

I just have this PermitSigner class that will sign exactly that message. And whenever we invoke this permit function, I'm invoking it to generate a valid signature, which you can do, because running the fuzzer, we know all of the other arguments. So we just generate a valid signature, pass that to the contract, and then our permits will go through, and then we will find the transfer from again that exhibits this donation-based attack that I've shown you. Also, there's another decision that we are making here. If you always generate valid signatures, we are never fuzzing the case where we are passing invalid signatures.

So there are always tradeoffs that you are making. And this is kind of what I want to do to sum up in the end. So we've seen a lot of things in this workshop, right? The very first thing, dump fuzzing does not work. Just generating random inputs does not work, at least for smart contracts.

It works for other use cases. The number of runs says nothing about the quality of your fuzzing campaign. You can run this for 500 days. If you're passing the wrong inputs and all of your invocations revert, you're not really exercising the contract. Fuzzing is not a black box.

There are no magic tools, not even in Web3. You have to always know what your tool is doing. You have to check what it's doing with the contract. Pay attention to that. There's a tight feedback loop, and you have to keep evolving it while you're building your fuzzing campaign.

Fuzzing also means making tradeoffs. You've seen this when I talked about the distributions. You have to make decisions from which intervals am I sampling. Am I missing anything, right? Am I missing any corner cases that I want my fuzzer to also cover?

At least if you do a fuzzing campaign for someone, make sure it's written down. Make sure it's in the code. Make sure it's in the audit reports that you're giving to people because those tradeoffs that you're making might be the corner cases that you're missing and that might be at some point exploited. Input distributions are tricky. I talked a bit about this.

You can do advanced stuff. You could use NumPy to make crazy input distributions to make them nicer, to make them more usable. Otherwise, you're back at making tradeoffs. That's just the way life is. The last thing I want to tell you is take fuzzing reports with a grain of salt.

You don't really know most of the times how much effort people have put into them, whether they just ran a fuzzer as a black box tool or whether they actually went and looked at what the fuzzer is doing with the contract. There's also a lot of stuff that we have not covered. This is a long workshop, and we have kind of stretched the surface only. We haven't covered passing time. A lot of reward-based attacks include some notion of putting money into the contract, then time goes by, then you expect some fair returns, and the exploit is getting unfair returns, so you have to pass time or block time in your fuzzer.

We haven't talked about pre-compiles. Then DEVM comes with crazy stuff, or Solidity comes with crazy stuff. There's receive and there are fallback functions, and together with fallback functions, there's the whole area of proxy contracts, delegate calls that complicates things even further. We haven't talked about contract upgrade attacks, so there's a lot of stuff to think about when you're fuzzing these things and just be aware of them. With that said, I just want to say thank you.

As I said, there's a step-by-step commit history in the GitHub repo. I will put the slides online. I will post a link to the GitHub repo if you want them. Here's my telegram. Just reach out.

With that, I say thank you, and let me know if you have any questions. That's a tricky one. Yes, so the question is, once you have a decent fuzzing suite built for a codebase, how much of the code is reusable for a future codebase? Depends. There are building blocks that you can reuse for the same flavor of contracts.

If you have vaults or I don't know, if you have a staking contract, there are similar things happening there. Especially in Ethereum, where there are a lot of standards, where there are a lot of EIPs that define the interface of things, you can reuse a lot of the basic building blocks. But you will always have to go into the project, look at the project, figure out what it's doing, and see what you are missing. For ecosystems other than Ethereum, things are not as much standardized. So there is a higher effort that you have to put into it, because even the basic contract interfaces might not be standardized.

Okay. Great.

Automatic transcript — names and jargon may be misspelled.