AI Audits - case study - 2,800 AI-Generated Tests, 22 Findings | Bartosz Barwikowski | ETHWarsaw [4]
ETH Warsaw·Sun, Nov 9, 2025, 12:00 AM
Speaker
Bartosz Barwikowski discussed the high reward in identifying bugs, vulnerabilities, and how he was able to save a company 10 million dollars. 🎥 Recorded at ETHWarsaw 2025 Follow ETHWarsaw on social media for the latest updates! X (Twitter): https://x.com/ETHWarsaw LinkedIn: https://www.linkedin.com/company/ethwarsaw Telegram chat: https://t.me/joinethwarsaw
Transcript
Okay, perfect. So, let's officially start. Hey, I'm Barto. Today, I'll be talking about uh a very interesting audit uh for which I use AI to generate 2,800 tests and it allowed to find 22 findings. It's especially interesting because usually AI is pretty bad out auditing, but this is very special case where it was actually the best possible uh solution.
So a little about me uh I only look young. I'm doing the hacking for around 15 years. I found quite a lot of issues and as I said last year I was talking here about the bug which was allowing to steal $1.1 million. I was very excited back then but right now I have already three such bugs on my account which uh and there's nothing I love more to finding another one.
Okay. So like uh in February we have a very interesting outage request. It was the audit for uh Quan virtual machine. And uh the important part about this is that this Quan project they were working on some quantum uh resistant blockchain and they had this idea that they want to run any kind of the application uh in their virtual machine. And to do that they basically uh were running uh Alpine Linux kernel and uh using virtualization.
So they can they could run an application as part of the blockchain. Uh but the problem with this is that you know uh if you want to have uh smart contract running on the blockchain it must be deterministic. So on every node the outcome of the execution must be always the same. And this is actually quite a problem. I will explain in a second why and what could go wrong.
But this was the task in general to find all of the sources of nondeterminism in their virtual machine in their virtualization. So yeah um there are a lot of sources of nondeterminism. So you know uh like you have different hardware different parameters memory speed CPU speeds and then a lot of nondeterministic functions and features in the system like threads memory clocks random number generations there a lot of them. So it's very important because uh you know as I said if the node cannot uh reproduce what the other node did then you cannot build a blockchain because you always need to have the same result after uh execution of the smart contract. So just let me show you how many of sources of nondeterminism there are just quickly because there are like more than 200 sources of nondeterminism basically everything uh what you have in your computer even like if your disk is having a different speed of reading the data this is a source of nondeterminism because on some computers it will be faster on other it will be slower the results won't be always the same so quite a lot and this QVM was using an open-source project from uh Facebook to help with nondeterminism.
This project is called Hermit. So actually a lot of this testing was uh testing of uh virtualization and this Hermit project where they have uh the all the tools for nondeterminism. So this was mostly the focus on the audit. Okay. So the scope of the audit was huge.
Uh we have those 200 plus uh sources of nondeterminism. It would be a nightmare to test it uh manually. We would need to hire a team of few auditors. We would need to spend a months to review every potential source of nondeterminism. They would need to pay like half a million for an audit.
So as you probably know they didn't have such budget. And the other option which I was thinking about was to use AI to check all those cases automatically. And usually it's hard to do audits with AI because in case of the smart contracts they uh are not trained on this. There are a lot of different type of the issues which are not known. They're very like project specific.
So uh it don't work very well. But this case was different because actually nondeterminism wasn't something new. We have basically those this list with more than 200 200 entries about every potential source of nondeterminism and this is not like a new knowledge. Uh so actually so AI uh knew how to kind of deal with this and it was not a new language uh because they wanted to run any program. it was possible to write just in a C or other language like C++, Rust, Python to test uh those uh issues and problems.
Yep. Okay. So uh one second. Okay. So before before deciding um uh about if to go with AI or not with this project um let me tell you uh what is my current experience with in general AI uh and the security.
So you know uh I started to as probably a lot of you if you are working cyber security trying to use AI to help with the audits and initially there were a lot of projects uh doing some audits there were bot races competition on code of for arena where you know you had those audits with AI and the best AI auditor was getting some rewards but um in the end it most of those project like didn't succeed and a lot of uh uh stuff with AI audits and reviews were cancelled or banned and actually this is currently one of the biggest nightmare of all the platforms with uh where you do this uh back bounties of crowdsource audits that they need to ban all the reports from AI because it's generating so many of them that uh it doesn't work well in the practice. Um it doesn't work well because there are plenty of issues which I will talk about in a moment. Okay. And here's the interesting story uh about uh me trying to cooperate with OpenAI on this uh project. So few years ago, OpenAI launched uh after the GBT4 something called OpenAI cyber security grant program and I uh submitted uh an idea to build AI 3D auditor for for the grant and interestingly we got to the second round uh to uh to the program and we had this amazing meeting with OpenAI team.
Everything was great. they were uh supposed to give us the credit and access but this was the last time we were able to talk with them because then they stopped answering any emails and didn't follow up. So this was pretty sad. Uh but it was due this time when there were when they had those issues with uh sman being kicked from being a seal and stuff and a lot of people get fired. So that's probably reason.
But one interesting thing is that look at this like two three weeks ago after two years they responded to me that sorry but your application didn't make it. So you know only took two years but thank you open AAI and yeah so uh there they starting this program again but I don't plan to apply but you know we have if you are doing something with open with AI and cyber security feel free maybe they will reply sooner okay so um let me explain you um what is the in general problem with AI and audits and coding uh which not many people talk Uh unfortunately a lot of people who post about AI stuff on Reddit and LinkedIn don't work with AI and they're usually just you know uh posting not something they tested or did themselves but those things usually don't work in practice. So uh we need to remember that AI is okay we can argue about this but uh don't argue with me just you know we can take the screenshot of argue with chat GPT on this sorry okay so yeah it's we need to remember this is pattern matching uh still it's not like intelligence it's hard for it to like understand what's going on like smart contract program to see every edge case uh there's a lot of problem with context um when it comes to AI because uh usually when you talk with AI and something on like a small prompt, let's say few thousand words, it works pretty good. But um when you start to having a longer conversation, you upload more code, uh it quickly the quality quickly goes down. And a lot of benchmarks don't show it.
They are they are working with uh small subsets uh or small uh problems. But um when you check the benchmarks which are testing uh things which are like 100,000 tokens which is like a complex smart contract the quality drops a lot and this is very important thing um for AI usually it's hard to it's okay uh in case of solidity AI usually is not teach how to deal with solidity and there are not enough examples so um it's hard for it to understand what's actually going on how to write good code this code which they have is often outdated. So you need to give AI all the instructions in the prompt but you end up quickly with very large prompt if you want to need to if you want to explain everything and then you basically need to explain every issue in the prompt. how to look for this and how to detect this and this usually is works only to very few very few limited issues and um that's and there's there's no something like training for a given uh programming language okay in you can train AI fine tune it but it's not usually what the people say and first you need to usually work with the best models to get the good results And then like training them is usually impossible. But even if it is, it's a huge cost and there's just even not enough data to train them for uh for working with solidity or uh auditing in general.
Um yeah uh that is quite a huge issue and and there's another thing which is very important for AI and testing. you know if you want AI to do something right correct you need to have a way to verify if it's going into a good direction so um the problem when it comes to the issues is sometimes it's hard to even define what an issue is for yeah I mean there are simple issue like let's say crash or like some authorization bypass but then you got a lot of more issues like for example reward stealing or getting some advantage or getting uh more votes than you should it's very hard to create the simple definition to cover everything. So this is uh quite a huge problem. And the last thing which a lot of people working with AI and myself as well are falling for is um the reproducibility trap with by this I mean that it's very easy to convince yourself that something is working uh correctly with AI and it's solving a hard problem which you have. uh this is because you know if you try very hard if you try a different approaches prompt instructions finally you will find a way to solve the given problem you have but it doesn't scale I mean if you try to apply the same logic with next 10 or 100 projects it won't work I mean for when I was doing stuff with solidity the simple example was that okay I I managed to find some issues like re-entrance or some economic issues but if I only change the order of the functions in the prompt it was working exactly the same just a different order of the functions it wasn't able to find those issues anymore and there are many traps like this so that's problem with AI in uh AI ideas in general okay but this case is different we are um trying to find all the sources of nondeterminism in the virtual machine and the thing is first of all it's a well researched field it's not like there are some unknown new sources of nondeter minism in Linux kernels or hardware of computer basically everything what can be found was already found detected and documented some somewhere at least 98% of cases which is enough and we can quick we can easily verify if the AI found a correct thing or not uh because we can just uh write test and see if the output the result of the test was not nondeterministic or not so this is another thing because we can define what the success this and then we have a standardized environment.
We have this virtual machine on Linux kernel on Quo virtualization. We can run it across multiple machines many times. So it's easy to build this testing infrastructure and we can also write with uh standard language legacy which is uh something that AI knows well about. So a perfect perfect case. Okay.
So uh our strategy was to have five-step plan. First of them was to go gather all the data about known sources of non- nondeterminism which uh in this case was basically using all the AI tools deep research and doing all the research about nondeterminism like checking every category telling AI to think deeper search more like check the posics documentation about the functions which of them can be nondeterministic or not check all the hardware sources of neterminist ISM everything we did like 100 deep researches but uh that was out of the data. So the step two uh so the step two was to organize it somehow. So remove the duplicates u organize it into categories and subcategories uh which we also did uh manually in this case. We ended up with like 270 subcategories of nondeterminism.
quite a lot. So, uh then we started to automate the next steps and I want to say that you know this project was uh co-created with my friend uh Yuri who was responsible for coding this uh Python implementation. They later will be a link to the repository because um this code is open source and you can find uh details about Yuri in the description. He gets a lot of credits for this project. Okay.
By the way, I met him last year on the Ethereum Warso and he didn't know anything about the web tree. Yep. Uh sometimes you know you meet interesting people here. Okay. So uh we started to write uh a code in the Python to to make it work and the first uh stage of this AI agent was a planner.
So we have all those uh categories subcategories. Now we need to have some plan how to deal with them. So for every category subcategory we define what we going to test in given category in subcategory and how we going to test it. So for each of them we want to have proposition of uh at least five tests which will be run and all of them should be at different. So after running this uh planning engine we end up with 76 categories and 240 subcategories and each of them had at least five test proposition uh what to test.
So yeah quite a lot of them. That's when 200 2,800 tests come from. Okay. So the next step was to actually execute the test those test. So uh that was the goal of execution agent.
So for every uh category, subcategory and test proposition, the execution agent was generating generating the code. Uh it was able to generate the code in any language, but uh we mostly did it in C because C is the easiest to test some very hardcore functions when you need to interact with some CPU like assembly instructions. So it was the best for this. So yeah, it was generating the code. It was executing this code across many nodes.
It was validating if the nondeterminism was found then the test was marked as successful and was moved to the next uh uh next uh phase next stage. Uh but if it wasn't if there was no if the test was deterministic then AI was given a few another tries to somehow improve this test something different to uh make sure that this nondeterminism is detected. So yeah, it was all conf configurable. So it was in the end up to configuration how many tries and tests you want to run. Okay.
So uh this execution agent created those 2,822 tests mostly in C. So yeah, there were quite a lot of them and we were using back then mostly deepseek and cloud set for writing the tests and it worked amazingly well. Usually it was able to uh create um valid programs in like two runs. So first it creates the pro uh program then we check with celang compiler if there are any issues if it compiles and if there were any issues uh it was getting a prompt to correct them and it was usually able to correct them pretty easily. And this is this important thing about uh why it was a good case because AI is you know trained on a lot of C programs.
So it knows how to uh run them and there's no like any innovation which it needs to test because you know 5 years ago the technology for those testing uh when it comes to C would be exactly the same. So we don't need to give AI any new instructions how to interact with this. It was already there. So this was perfect. Okay.
So yeah um we have the test and around 4 400 of those tests were uh resulted in nondeterminism or some issues and that's quite of a lot of tests. So another like sounds like another month just to verify them and what is the source of nondeterminism uh in in those tests. So we had to create uh another stage which was review agent and in this stage um you know uh the AI agent would basically have running those tests and trying to find out what is the source of nondeterminism in this test by removing some code and checking if the test uh still fails or not and then trying to find this exact one one or two lines of the code which are responsible for the problem and after And after running this uh it was able to categorize everything into uh 22 separate issues and we had 11 issues which were a problems with hermit this uh open-source uh execution environment from Facebook and those were crashes and panics and aborts which was pretty surprising for me because actually it's it was amazing that you know this approach was able to find a lot of issues in other project which we weren't even trying to find because we were looking for nondeterminis but it was finding a crashes and real book bugs so it was like fuzzing with AI uh so that's was uh pretty surprising and it also found 11 uh sources of nondeterminism like uh yeah there were a lot of them so yeah I mean you really need to be a very a a geek to I think if you want to listen about them so let's skip about this okay so yeah we created a final report which is available on haken website if you go to hacken uh audits quan uh and you will enter quan you will find the report where you can read about all the issues which were found and what was the problem so yeah 22 findings and it was all done in around two months where we spent like one month on working on this AI stuff. Other time uh we are spending just on manual audit because it wasn't only AI audit. I was also checking a lot of uh things uh manually uh just to be sure and then um they were able to fix all the issues which is quite uh a big accomplishment.
So yeah, after a month of uh time, they came back with all the issues patched. And this is another great uh showcase why this approach was pretty good because we just rerun all those tests and this tool again on a new fixed code base. We didn't need to like verify everything by hand. We just rerun what we had and it all worked. it didn't find any new issues which was uh an impressive job done by the quant team.
Okay. So yeah what we managed to achieve with this approach. So first of all we saved a lot of time you know instead of half a million audit uh we made it for much less and it's not like we don't like money I guess uh they just wouldn't be able to afford this and we get a superior coverage. Uh so a lot of things were tested. Uh basically every like posic function every source of neterminis even some hardcore function like uh vector operation or assembly functions or some built-in CPU randomness uh was found and tested.
So that was amazing and it would be very hard to do manually. Uh yeah and this tool was uh reusable and future proof. So you know whenever you change something in your implementation of the virtual machine add new features you can always rerun it again to try uh to find out maybe you made made some mistake and now it will be able to find it. Uh so yeah this is uh a perfect thing because you know write once and use kind of forever. Okay.
So, um that was it when it comes to this audit. But let me probably out of you wonder uh how the future looks like when it comes to AI and the audits. So in general from my experience uh AI uh won't replace like the human auditors uh because uh when it comes to working with like a new stuff uh which was just created and there are not very examples there are a lot of edge cases and like lot of interactions and things to account for it's very extremely hard for AI to do anything with this that's why we don't have any good projects right now in general I have a lot problems with memory. This is like the biggest problem with AI right now that it you cannot really teach it something new and it doesn't really remember what you were doing with it before. That's why we don't have any AI agents who work like a personal assistant who knows you and can like recommend what you are doing.
And if they are, they are usually very limited and they quickly start to make many mistakes after uh some time like a like a month or week. So yeah uh so AI is more like a tool which just will speed up all the work maybe testing maybe confirming uh some issues but uh it's for sure not immediate risk for uh for the senior developers and senior auditors I mean we need some breakthrough in AI uh memory to let's say be in danger when it comes to this. So you know when you when um AI assistants will start to be a norm and they will be very good this is the moment where you need to start to be afraid if you are working in this field so this is how I see it okay that's all thank you for listening I hope you enjoyed and we even have few minutes for a questions do we
perfect first question I will check if I have this link to repository somewhere. Yeah.
Hey,
doesn't matter. So my
Oh no, no, it's working now. Uh yeah thanks for the presentation. Um my question is in these steps uh like what was the orchestration tooling that you were using to do this or maybe it doesn't matter. Uh second question is like what kind of agents you were using? So nowadays you have this code and codecs which maybe you didn't have before.
So for the agentic behavior what did you use? Was it in house? And third question is a human in the loop like how much human was involved in each of these steps and at which stages? Uh if you can clarify those parts.
Okay, sure. So um when it comes to the agent uh we were using I don't remember the name but you know back then there was no cloud code. Uh we were using some uh different tool to write this uh code. Come on I can I don't either. Yeah, we are using the Ader implementation to write the code back then but it wasn't like necessary.
uh we're just you know have we were having few defined steps uh of interactions with AI like okay first you write the code then we run this code and if there are any issues with this code we just send you back the issues and now write the new code which will correct this and for the other step like with the categories and subcategories we are usually using just a single prompt for like two or or sometimes more like uh the first prompt is okay propose the categories then criticize it like think what could be done better and then implement the new version and it was just done by using structure output. So it's not like there were some advanced agents uh used in this just interactions with AI mostly with this deepseek and with later with uh cloud sonet which worked pretty well but right now if I had to do it I would probably use cloud code because uh it would be cheaper uh but it would take way longer that's the problem because you know it would take like probably two weeks with a cloud code to run all the tests. So, but it would be probably much better. H it depend it will depend on the budget. Uh if there were other question you need to repeat them.
Okay.
Awesome. I think we have time for one more. You were you were next.
Uh I was wondering when you were uh using AI for the fuzzing. Um uh you mentioned it briefly but did you run into anything where it was hallucinating returning different results each time you went through? Did you have a set number of test? Did you run it a set number of times to try to reduce that or how much of that part of that was?
Okay, so of course you know and a lot of times AI was doing stupid mistakes and something wrong. But here we had this advantage that we are able to verify if it works and if it's doing the good job because first uh is writing the C program and this program needs to be to compile and run. So this can be verified. it doesn't compile then you know we tell it about the AI and it gets another try and then we have the results of the execution of the program so uh we can get this output of the execution this std out and we can see if it caused nondeterminism or didn't so you know if AI was starting to hallucinate or doing something wrong and uh the tests weren't compiling and were passing then it was informed that it's doing things incorrectly so that's why it works in this Okay, I think one last one minute
you
I will answer you.
Okay, thanks for your presentation. I want to know that uh how to use AI agent to um uh to audit the onchain system but not as single smart contract which which means that uh consider onchain system you must to um deploy on the your AI agent on every node of the uh of the whole network. uh and then and and may contain some some interaction with with this agent on on each node uh so that to to an analyze to to get a data and analyze uh network traffic or something else how to do that
I don't know I don't have experience with something like that and in general when it comes to blockchain and AI on chain and some trading agents I know nothing about this.
Oh, sorry about that.
So, yeah.
Anyways, and thanks for your presentation and I I think I would like to get some answers maybe um from anyone here who has experience or some or external experts. Okay, I got it.
Awesome. Thank you so much. Thank you.
Automatic transcript — names and jargon may be misspelled.