Alexandru Andrei Dobra - The Agentic Layer of Smart Contract Security
ETHCluj Meetup·Wed, Sep 9, 2026, 12:00 AM
Discover how AI agents are becoming a new, powerful, and affordable layer of smart contract security, with insights from Nethermind’s AuditAgent and real-world AgentArena bug bounty competitions for protocols like Lido and Uniswap.
Transcript
Okay, thank you everyone. Uh my name is Andre Dobra. I'm a full stack engineer at uh Nethermind and today I'm going to talk to you about our contributions to the agentic layer of smart contract security. Uh, as you might know, the attacks that we're seeing in the in the ecosystem are getting more uh are happening more often and uh um yeah uh and they are happening more more often and we've seen just in the past month over $600 million being uh stolen from uh from some of the major protocols out there. And so it's becoming clear that this problem is um is occurring more often and the attack surfaces are just multiplying and they become very easy to to uh these uh these exploits become very easy to uh make because um everybody has access to these agents.
Um and now uh it's very easy for everybody to just find even uh even hints at an issue and then to investigate further to create the test cases. And so everybody's getting these tools that can be used both for good and for bad at the same time. And so uh in this age of of AI security tooling you have to become proactive and uh you know how it is with security audits is good to have as many as possible. It's not that you said okay I got the security audit from this company and then from this company and now I'm covered. No in uh today you have to be proactive.
You have to start using these tools that help you throughout the development process. So you get your code base to a state that is already pretty good at the audit uh state, the formal uh verification state and so on. Uh because earlier earlier in the process, you might uh lose track of certain uh little gimmicks or trade-offs that you've introduced in the protocol. And then if you have these agents being uh pretty much an assistant uh a co-pilot for you or whatever you might want to call it, uh they assist you in finding these lowhanging fruits, low hanging bugs uh throughout the development process. And then when you get the audit, you can have the auditors focus more on the serious issues.
So more on those business logic flows that you haven't seen before. And um basically this uh this agentic layer can comprise a bunch of tools. So we can think even about uh or a lot about uh solutions such such as uh monitoring onchain monitoring of activities because it's no longer about just finding the issues in solidity because yeah this has been going on for like we've been finding them uh for a long time. We know about these issues. We have uh we recognize many of the patterns but today many of the hacks are coming out of of uh lacking uh in security when it comes to infrastructure and to process right what has this uh protocol established as a process from themselves how uh are they using a multisig is that multisig respecting a certain uh a standard are are those people taking all of the proper security measures and when they get hacked do you have signals for that do you have the proper uh flags that uh give you enough time to react to them.
But that is about the broader broader perspective. And now I'm going to talk to you about this our specific take on uh on this uh on this agentic layer of the smart contract security. And so one of the first uh the first products that I'm going to talk you about and maybe it's the more obvious one that most security companies end up developing at this stage. Uh so this is our audit agent and you can see there one of the there's this uh quote from from uh Twitter from Vitalic uh where he's saying that one of the applications one of the applications of AI that he's excited about is this AI assisted formal verification uh of code and bug finding. So these use cases have been identified for a long time.
I believe this is already two or three years old but it's it's becoming more and more valid that these tools are here to stay. So it's not a phase. uh imagine that uh these tools are becoming essential for our auditors as well. So they use it in their process. Even if the AI doesn't find necessarily the core issue behind it can still give the auditors a hint about okay what is not respected here in regards to the documentation in regards to the initial specs that they decided for or are they even considering at this point the tradeoffs that they accepted in the past because this is something that we see often.
So a protocol starts developing they get the audit they accept certain tradeoffs or behaviors and then later on start expanding and building on top of it and these trade-offs are forgotten. So it's very important to build out this context uh that you're uh for the whole protocol that you're building. So what is uh audit agent? Audit agent is our AI powered uh smart contract auditor which we see a lot as a pre- audit tool. So you have this companion throughout your development process which you can integrate in your CI.
So whenever you make some updates, you get some uh some insights about what might be wrong in there. You get the audit reports uh in hours, not weeks. So depending on the complexity of the codebase, it could take, you know, maybe if it's something under a thousand lines of code and quite simple, it could be an hour or two and then up to five hours depending uh depending on the size. Uh and it is able to identify critical vulnerabilities. Uh we've seen this in practice multiple times and so it's very valuable to identify these issues earlier.
Um and then the these are the core principles of it. This is to a large extent how most of these agents uh in the industry work because you have first the contextual analysis. So what and why and this is where the where the context is very very important. Uh and this is something that I stress a lot when I have these calls with uh with clients and possible clients that um want to use the tool because it is a self-service tool. So you can just go ahead and use it right now.
you uh upload your uh you select your repository because it's integrated with GitHub. Uh you select the repository, you add the docs and this is the part that I stress out the most the importance of the docs and not and here it's not really important the size of it because from a certain point onwards uh the process breaks down because as we know when the context goes above let's say 100,000 150,000 tokens for most of these uh LLMs then the quality starts breaking down dramatically. So we can reach even a 50% efficiency if the cost if the context is too big. So we stress this thing about creating a kind of structured documentation. Um and here we're talking about uh the interactions that you're expecting the specs but not the full description maybe as you would deliver to a human developer that is joining the the joining the project but rather as a structured schema that can be a lot more useful to the LLM to understand your codebase.
Um so you have this contextual analysis what are you building and why are you building and what are the implications what are the trade-offs then you have the structural analysis so how have you actually implemented this so then based on all of this context that you've provided it goes ahead and tries to understand what you've built there and if you haven't mentioned for example what kind of um roles which what are their um what they can do so let's say you have an admin uh admin role that has a lot of power over the over the protocol but you haven't docu documented one of these uh functions that it can call then the uh the agent can go ahead and flag it and say okay this doesn't to seem to be fine. It doesn't respect the the specs that you've built and then it creates basically a false positive. It might look like a false positives but guess what you actually forgot to mention those things in the documentation and in what uh in in those specs. Um so in itself it can flag issues both in your codebase and in your docs. So it basically focuses on the inconsistencies between the two.
Uh and so we get to this point where uh previously developers really didn't want to write the docs like that was the most boring part that you want to get to. But now it's becoming essential for the whole uh for the whole uh process of developing when you get these AI agents into the mix mix. So you could start thinking about um like you have testdriven development but you what if you now have a kind of document-driven development. So throughout the process, one of the most essential parts is documenting every single bit and documenting every trade-off that you've accepted and so on. So s such that for the future for this uh consumption by agents you are prepared and you are ready to to deliver the necessary information.
So there's a structural analysis then you have the multi- aent system which is actually audit agent in itself is a multi- aent system because there are a bunch of different strategies uh used a bunch of uh different prompts of and angles that it explores. So you could end up with having maybe more than 100 calls to to to these LM providers in one single analysis if the code base is really large. It also has access to a bunch of different tools like static analyzers, web search, uh testing and so on. So it it uses all of this to either validate or to to build up the build out the the vulnerabilities in the report. Now I would like to show you some of the progress of audit agent.
And here we can see how it was basically in January uh last year. And we had sample uh for this testing of 98 vulnerabilities from five different repositories and at the time so in January last year we found only 12 out of those 98 vulnerabilities right and then you have total findings 87. Uh so so at this point we were finding just 12.2% of the of the vulnerabilities in those reports. Uh and so this was basically in its infancy.
At that point it was only 6 months into development. And then we go to the next milestone which is just 6 months later. So we went from January to June last year. And then here we got to the point of finding 51% of all of those uh all of those findings. But the the the other critical part here and the other thing that we're following when building such agents is how many false positives were there.
So for those 50 valid findings out of the 98, we also had a lot of um quality assurance findings. So we had 94 here, which adds to it, but at least they're flagged as info, so you don't have to necessarily go through all of them. But then we also had 196 false positives. So you get 50 correct and 196 false positives. And this results got us, you know, to rethink about the whole process because you can't have users dig to so many false positives in a report uh because it just makes it a chore.
and they might then no longer focus on the the actual good things. So then we took a big hit and we basically went down to 35 out of 98 but we also started reducing massively the false positive. So we dropped them by about 25% and then we started focusing on both things at the same time both the detection part and the validation in order to provide that high quality of a report. So then even six months uh or about seven months later which is in March this year we now got to 60% 60% of all of those vulnerabilities are now found by audit agent. We've reduced the false positives from about 200 to 69.
So now you get a report that is almost like 50 50% false positives and uh 50% true positives. uh but do keep in mind that of out of those false positives as I mentioned earlier they might not they uh might not uh flag um full issue let's say that is exploitable but still it can flag an inconsistency between your docs and your codebase. So then this is something that you might have to work on further. So this is where we are at uh right now and now I would like to show you a comparison between uh uh audit agent and the top tier models. This is uh the comparison from this year because EVM bench came out and we also u ran these these benchmarks to see how it performs.
Um and this is a comparison from last year. So I would I would go start from here. So this is a comparison from last year I believe in June or July and audit agent or or a bit a bit later actually. uh audit agent was finding 48 out of those 98 vulnerabilities that were in our internal benchmark and Opus 4 at the time was finding just 17. But do keep in mind that these results were coming out of our specialized prompt engineering.
So it wasn't just the models being told okay find vulnerabilities and that's it. And now when we go to the EVM bench that uh so this result is just one month old. We see that audit agent finds 67% out of all the vulnerabilities uh in the EVM bench uh bench in in this benchmark and then we have claude oppus 4.6 finding 47%. And now when you look at this graph, you might be thinking okay this gap is closing actually.
So like what uh why are you showing us this like why why would you go on and flaunt something like okay last year you were finding three times more than the next best frontier model and now you're finding just 20% uh 20% more uh or 30% actually compared to to oppus. But the difference here comes from the what are these vulnerabilities that we're finding now? What were we finding last year? They think were those were very complex vulnerabilities. No, those were like the very lowhanging fruits that were quite obvious that a trained auditor can go ahead and find in maybe a few minutes or or an hour or something in any case uh very in any case a very short time frame.
And now the battle for conquering the next two 3% is becoming even much even much harder because for each new new percent you have to identify even better strategies. You have to get the agent to build more complex uh chains of thought so that it gets from uh from a hint that there might be an issue there to actually building a complete exploit and identifying the the full vulnerability end to end. So these percentages actually uh are much harder to gain at this point. So now the gap is not actually closing but finally the frontier models are catching up and they are finding those more obvious issues more easily. But the battle still stands in the top percentages.
How do we get to the next 2%, the next 3%? And here it might require exponentially more development. Uh, as I said, audit agent is live. You can just go ahead and try it for free. You can scan 500 lines of code out of the box.
If you complete like a small survey, which is just a few questions, you can get to 1,000 lines. And actually if you if you would like to learn more about it or you're curious to test it, you can talk to me after afterwards and I can arrange to have a coupon for you so you can go ahead and use the best uh the best scan quality that we have and see what you can find in your codebase. And I'm really interested in your feedback. So yeah, if you would like to test it out in full force, you can uh talk to me afterwards. Um now I would like to talk a bit about the uh landscape that is developing around us.
So we're having growing AI adoption. We're having shorter release cycles. And we're seeing this everywhere. Most companies are pushing harder and fast and and and more and more in this direction of okay, now we have AI. That means we can do a lot more.
So now we want to push things out even uh even faster. So you have this shorter release cycles. This means increased demand for automation. And this increased demand for automation also implies increased complexity of AI generated code code or spaghetti. That's why I added this image there because that's the kind of spaghetti code that you have to dig through now in the review processes because review is becoming the bottleneck and you know you're not going to catch me review that kind of stuff.
So then you end up adding a lot more agents into the pipeline to test everything out to find the more obvious issues such that you get to a clean uh clean code base uh or a cleaner uh code base. Um so then we have this increased complexity. It means we have an increased demand for code auditing and for AI review. And then this AIdriven development introduces obviously uh novel exploits or not only novel but rather silly. So we've seen such cases where we see things going into production that have been barely reviewed and this implies more bugs, more issues, more exploits and more hacks.
And we're seeing this across the board. Um and then you see that these tools are kind of isolated and these isolated solutions um are not quite hitting the mark. Um because these agents yes they are getting better and better but they are not yet at the stage of replacing a human auditor or some a person that deeply knows the the protocol and then can find the issues in there. So we have all of these separate separate tools and we're thinking okay what if we put all of them together actually so we end up with this need for orchestration and putting all of these different solutions together such that we create a more more powerful tool and that's how we actually got to to the next uh thing that we're building which is agent arena and this came out of observations through of how the market has been developing and how many of these security agents are coming out. So then if you have one agent which can find a certain set of issues and then you have another agent that is more specialized let's say on me extraction or you find different kinds of agents which have a certain specialty that allows them to find a specific type of of issues.
So then what if you would put all of them together? So you end up creating this this kind of multi- aent system that is driven by independent builders because what agent is is a space uh for competition and for collaboration as in itself it is a an audit competitions platform for AI security agents. So this is the place where you would uh if you are a builder, if you would like to get into this domain, you could start building your own security agent. you plug it into the platform and then whenever a new competition comes up and your agent finds uh a valid vulnerability then it gets paid out uh like in u bug bounty platforms uh for for auditors and we've seen this recent sad news that uh I think two days ago that code arena is shutting down and for uh for security auditors this is quite a sad day because code arena left a mark on the whole auditing uh auditing space uh and So even though kodarina is shutting down, we're coming in with agent arena. So we're trying to get into this uh this uh new world of of development that is AI powered and well AI we're seeing that it's uh it's here to stay for the long run.
Uh so we're building these tools then and these these platforms where we can aggregate their effort. So further agent arena is open source and uh and transparent and we have this uh uh this is regarding the arbiter agent because the platform in itself has to has to remain uh um transparent when it comes to the arbitration and so we are planning to run this arbiter agent in the future in the in a trusted execution environment as well such that the process can be uh reproducible and can be um tracked by people and then we have this transparent onchain uh distribution of the bounty. So we have these contracts where you can see the the bounty pool goes inside and then the agents get paid out of there. And now the functionality in short of the arena is the sp the the sponsors those that want their protocol to be audited they send their code to us you know they get it from GitHub or whatever they add the documentation and so on. Um the competition starts then the agents fetch all of this data the repository in the docs they analyze it and then they deliver the results.
um and the arbiter judges and then optionally we can have a human auditor as well to validate uh validate the results um because at this point if you're just validating AI with AI well you can get a certain quality of the final um the result set but it's definitely not not something guaranteed. So you would like to have a human auditor there to give you the final verdict and then in the end after we see the results we can also distribute the rewards. Now regarding the arbiter agent and how the competition process works, we have a very short submission window. So these agents have just 24 hours to download the code, run the analysis and submit the results because the key here was that we do not want human auditors to participate. So they do not uh influence the results in any way.
So for example, you know, if we gave a week and an auditor was building their agent and they, you know, besides running the agent, they start digging into the repository and finding some issues. If they submit that and that gets validated then their agent would become would go on top of the leaderboard and well that's not a fair result. Uh so we came up with this 24 hours submission window. Uh and then for the after the agent submit the results you get to the arbiter agent in itself and it this one has three steps. It did it identifies the duplicate findings within the same submission and across the set of results groups everything.
Then you get the validation part. So it assesses the validity and quality of these vulnerabilities and discarding the highly probable false positives. And then you get to to reporting. So it assembles this final set uh set of results and uh sends it to the main platform uh where it can get either human reviewed or directly uh the rewards can be distributed and as I said earlier we plan to run this in a trusted execution environment in order to ensure this uh uh temperroof evaluation of of results. Um as I mentioned the key features open and transparent validation fair distribution uh an open ecosystem so this is a key uh feature here everybody can join we're not whitelisting we're not blacklisting we're just letting you participate now we will follow the behavior of certain agents so if we see that they are spamming us or uh certain uh kind of malicious behaviors then yeah we can uh act on it and we can contact the the developer or just uh take some measures but in itself the ecosystem is open.
So if you are just very good at prompting, if you're building agents, even if you don't have the security background, but you know how to build a AI agents, then you can go ahead and join the platform uh and also start u getting some rewards because this is one of the one of the things that I I quite like about agent arena especially is this part about democratizing uh first the security auditing space, but also we're creating this space where everybody can participate in the new agentic economy. So you have all of you have a lot of people that are now thinking know that they're losing jobs but uh they can already start building these agents and make and getting rewards out of them. So what if we start creating this collaboration places where everybody can start pitching in and uh reaping rewards out of this new agentic economy that is being developed. Um then the next uh the next uh key features the real-time contest results that you you get very fast. Uh then you have this broad com coverage carving out of the multiple specialized agents and in itself the platform uh encourages this innovation in AI based security.
So we would like to see more and more builders joining and then uh building uh together towards this uh better better space for uh security of smart contracts and not only uh this is what the platform looks like right now. This is just what you would see as an agent builder. You see your dashboard you can have quick access to the docs to connect your um your agent. You can see you also have a test task here which you can uh which you can run to see how the agent performs. you connect you know your uh you have your settings over there you connect your wallet for the rewards and uh and so on.
Now we also provide a security agent template. So this is a very basic AI security agent that we're providing to everybody open source. So you can just go ahead and uh start building on top of it. And in itself we we were just running it sometimes in competitions and it has managed to find one high and I believe a few medium vulnerabilities and this is a very basic agent like it just fetches code. It assembles a prompt for finding vulnerabilities queries the AI and then delivers everything under the correct format.
And this very basic agent that is actually just one prompt managed to find to find the high and medium security vulnerabilities. This is open source like you can just go ahead and use it. Um so you can you know clone the repository do the setup you can test it both in local mode and in uh and in server mode and then you can follow the steps in our docs to connect it to u I mean actually it will be already connected to the platform. You just have to do some small configurations and you can scan that QR code to get to the template itself. Um now another highlight that I wanted to give you is this um this collaboration that we've done with Lido.
So we've run several competitions for them. And as you can imagine, Lido has some of the most u uh audited contracts in the space. They have uh basically over 20 billion dollars in uh staked uh ETH in their platform. So we're talking about a very uh mature codebase overall and a protocol that requires high standards. So at their level just getting human uh human auditors to see the codebase is not enough especially since we're seeing so many hacks.
So it becomes mandatory for such protocols to start looking for other solutions that complete these uh uh this overall um layer of uh of security. So you no longer have just the manual auditors but you add on top uh either monitoring, you add on top audit agents, you add on top these competitions from agent arena. So you try to get as much security as possible and as many layers and what we have observed in these competitions there were three competitions where we found a total of six medium severity vulnerabilities. Now do keep in mind that some of that code was uh uh not fully audited. So we're talking about different stages of development but uh still seeing that the agents found uh these medium severity vulnerabilities and six of them in in such a codebase it signal to us that okay there is value here these agents are able to find meaningful vulnerabilities even in more battle tested uh code bases and now the current plans what we're doing now is we're onboarding more agent builders more sponsors um and we're improving the arbiter agent Because ideally we would like to get this platform to a point of uh complete self-service and without any other uh uh human contribution.
So you would no longer have that u human auditor validating the results at the end. Uh and therefore we are working uh on improving the arbiter agent and then further for June July we want to further maximize this transparency. So what I've told you about uh either the process of the arbiter starting to release results. Now you know releasing results is um cannot be done freely because it depend on uh on the input of the sponsors. They might not want to release the full set of results but we will try to put you know as many of these results out there as possible.
So everybody can trace what where these agents found the vulnerabilities and how many they were. Um so yeah this is the maximizing the transparency then streamlining the competitions and automating the full reward distribution and then in the long term obviously we're thinking about full autonomous operation community governance and um this ecosystem of specialized agents but also we're thinking about a tier system because you might get more and more agents so then you're getting spammed. So what if you create this kind of two tiers for the competition where the first one is the proving ground where you where we can see if you're actually able to deliver some meaningful results and then we have this other uh select few set of uh more uh of better o auditing agents and then uh these ones could go to some specialized uh competitions. Um all right and here you have my uh my contacts and links worth uh agent arena my contact there you have the contact of nethermind security overall and uh this is another group for the audit agent that you've seen in the beginning. Yeah that's about it.
Thank you.
[applause]
Automatic transcript — names and jargon may be misspelled.