New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Building Trustless Onchain AI: Verifiable Federated Learning | Andrea Rizzini - Horizen Labs

Ethereum DenverMon, Mar 9, 2026, 12:00 AM

Autonomous AI agents need to learn from private data without sharing it. In this talk we explore Verifiable Federated Learning (VFL): using ZK and cryptography on Ethereum to enable trustless, collaborative model training and on-chain aggregation for decentralized AI fleets.

Transcript

All right. And now it's time we move forward towards more AI and agentic stuff. I know a lot of wipe coders are still with us here in the audience. And it's time that I invite on stage Andrea who's a researcher at Horizon Labs to talk more about building trustless onchain AI. It's all the buzz and Andrea is going to walk us through it.

Andrea, are you ready?

Let's do it. All the best.

Hello everyone. So, so today we are going to talk about uh federated learning and uh uh how verifiability can uh fit on top of on top of that and uh we'll see how verifiability can enables a lot of advantages in federated learning in terms of security and in terms of many other things that uh we are going to go So first of all a few mentions to my colleagues Franceski Marquez and Tomas Gallardoni who are colleagues both from the university and the and resin labs. So let's uh let's start and uh let's first of all take a look on motivations and on what uh federated learning is as a first thing. So let's say in the federated learning framework the row federated learning framework we have many clients that trains a model locally on their private data. So we have the advantage that the data set to remain concealed and uh these local models are then used are then sent to an aggregator that aggregate the model and then it um redistributes it to the to the clients and there is this sort of iterative procedure.

This is mostly used in healthcare in finance and in those fields where the data set is uh is containing sensitive informations. Erh but uh of course for instance in healthare we can we can we can say we can assume that hospitals are trusted but now the federated learning can be used also for many other use cases and applications. So in the general case we don't have any specific guarantee that the the entities are trusted. So these works rely on that. So let's take a look quickly at the agenda.

So we'll see some preliminaries on federated learning uh something on verifiability. Then uh I'm going to explain you a framework that uh we built uh to enable verifiability on federated learning. Then we'll see the state-of-the-art solutions and then we'll see what where the research should focus and then of course we we will talk about future and u AI agents because of course is the most hyped stuff of the moment. Okay. So uh let's take a look first of all at the federated learning at the federated learning.

So h I know that you are both technical and and not technical person. So I tried to reach out uh both of you. So for the technical person focus on the procedure for the not technical one focus on the image. So uh as I told you it's an iterative procedures where you can see many clients that train that trains a model locally. Then this model is sent to a an entity called the aggregator that uh takes these partial models aggregate them and redistribute them to the clients and then if a stopping uh criterion is met that's the final model that will be deployed otherwise proceed with uh with another step with another iteration.

Erh so uh there are many advantages uh of course you can obtain a better model since uh you use more data from many clients so you can obtain a model that generalize better but uh this comes also with some drawbacks for instance the fact that there are many entities can increase the attack surface uh we can find many attacks related to federated learning. There are inference attacks that are inference and gradient inversion attacks that are nicely mitigated let's say just by the fact of sending the weights masked in some way. So there are secure aggregation and differential privacy stuff that you can use. Then there is model poisoning that's mitigated but not so mitigated. Uh then we can find global model tampering selective dropping so the aggregator that uh get rid of some of the updates and free riding the fact that uh clients can uh just participate but without collaborating.

So uh trying to get the model without uh the actual training. So smarter mitigations and mitigations for many of these attacks can come from verifiability. So indeed verifiability can really help in mitigating certain type of attacks. Erh so here just a quick overview of uh verifiability itself. So uh an informal definition is uh uh that verifiability can enable the possibility to a third party of verifying the computation without uh the need of reexecuting it completely.

we can find two forms of verifiability zk based and tb based and uh each one of those comes with advantages and drawbacks. Erh for instance ZK based are based on cryptographic assumptions. YT based are based on hardware assumptions. Er, of course you you you will you will certainly have heard of snark and starks and u for uh DK and uh you will have heard of uh Intel SDGX and Intel TDX and AR zone 40 based. So this just to say that there could be there are multiple ways to achieve verifiability with advantages and drawbacks.

Erh so let's take a look at the uh verifiable framework. H there are many challenges in machine learning uh d that inclusion or exclusion could be one thing that we want to prove. er provenence and authorship of the data set and um one specifically for federated learning is collaborative correctness over partitions. So verifiability can manifest on the verifiable data usage on the training part and on the and then and on the aggregation side. So here this slide I know it's quite dense but uh just to show you that uh we we have built this uh framework uh where it's possible to make uh those seven claims approvable and verifiables the most important one are data binding and training correctness that are related to the training and AG1 and AG2 that are related to the aggregation.

So just to take a quick look at the algorithm, you can commit to the data and you can use you can essentially prove the the training procedure taking as input the the committed data and getting as output an endle to this um to the weights that can be they can be encrypted or they can be in clear and we can commit and push everything on a transcript that that can be in this case a blockchain. Uh for the aggregation side, this algorithm can make verifiable AG1 and AG2 uh where simply the aggregator takes information. So the weights from the blockchain verifies all the proofs that we call evidence because we are generalizing we are trying to include both GK and quotes by it is so where you see any I'm intending an evidence so if all the evidences are verified so if the the outcome is one to the verify function uh then the agree Aggregator can use this admission set and aggregate all the valid and proved weights obtaining the the global model that then will be appended on the on the transcript and uh and then uh it can start the next iteration where also the client can prove the the handle can verify the the evidence for the agree. ation this is the state-of-the-art so many works are I mean many researchers and many university are working on that they are 13 papers uh really recently really recent from 2023 to 2025 so it's really at the state-of-the-art all of that and this is a first slide on the comparative analysis uh you can see I've divided that table between works relying on ZKP and works relying on TE's and uh by the graph you can see that that the majority of the works are focusing on the aggregation part so on a G2 so the correct aggregation that's I mean that's why the that's because the aggregation are mainly linear operations that can can be circuitized really well. Uh while for instance you can see that on the training part there is a lack in research because the proving scheme the the actual proving schemes are not ready because they are not so how can I say uh they are not so good for nonlinear operations that happen on um on the on the layers of a of a model being trained.

Okay, then I put also the second slide on comparative analysis showing on the first table that many works relies on the blockchain to verify proofs actually six out of seven and um uh there is just there is actually uh only one work er relying on the on the for the aggregation. So the blockchain can really play a crucial role for as a verification medium but also as an auditable transcript where to record all the process. Another notable thing is that uh implementing verifiability on top of all of that can mitigate certain type of attacks. Erh and that combined with privacy preserving techniques can really uh we can really get um a solid and sound uh framework. Okay.

Uh these are benchmarks. So cost analysis. Okay. So claim one do you remember is the one related to data binding. So proving that I'm using certain data that can be implemented simply with a commitment plus Merkel proof and it's relatively cheap because you know a Merkel proof is cheap and you can also see that most of most of the overhead comes from claim two and claim three.

So claims related to the training uh that um from the fact as I was already mentioning that the the layers are nonlinear and uh that's not circuit friendly let's say while for the aggregation many works implement it h it it absorbs most of the overhead because we have to to make a weighted average of the all the uh weights coming from certain time hundreds of parameters but it's easy to implement in a snark because uh basically they are just a linear operations okay so where should the research focus we state that many claims that many works don't implement are feasible and easy to implement. So we really suggest to researcher to focus on that and then um yeah of course focus on new proving systems to circuitize and prove all the layers of the training. we are really behind on that and then a cool thing would be to explore this um coark that's a new paradigm that that's a sort of MPC where many clients can generate in a collaborative manner a proof and this could possibly remove the aggregator uh notably if uh this framework will be implemented. We will have many proofs, many evidences that are both ZK proofs and T quotes to be verified somewhere. So how to verify them and where to verify them?

Well, of course there are there are layer twos layer two like uh optimisum and base that are already safe and cheap options. Erh when uh when we are talking of uh snarks erh so you can see for instance that for a for a 16 proof that uh for a 16 proof with two public inputs this take fraction of sense in uh most of the layer 2 but uh we are developing zk verify that's this blockchain tailored for um verifying proofs efficiently and uh cheaply. And you can see that for the same proof. So for a G 16 proof that cost 220K, we can basically add another zero in the 0.00 stuff that you that you see there.

and you can decrease even more the price with respect to the h to the cheapest option in the EVM environment. Okay. So, uh to conclude, we haven't talked about agents. Let's see how our framework can be integrated also with agents. So, let's frame it like that.

So let's imagine a fleet of agents where each agent where each agent has a B has two uh let's say two model one LLM one is an LLM of course to decide actions to to orchestrate the of course the the yeah the the next action so the actions of the agent and another specified for a specific task uh and the second model of course is is the same across all the agents and this model can be can be trained in a federated setting. So let's imagine different agents having different data training shared model. uh this can potentially I mean they can potentially achieve a a better model because of this uh variety of data that generaliz generalizes better on the task. Erh so we said that uh we are not trusting entities that u that are me that are human based let's say so how can we trust agent how can agent trust each other well we can rely on the VFL framework of course uh we can have two architectures the first one is the classic with uh many agents acting as client training locally the the model and they can rely they can push stuff on the VFL uh medium so to the transcript and the the aggregating the aggregating agent can fetch stuff from the transcript verifying correctness of the training performing the aggregation and push stuff on the transcript. So the exact same thing that we have seen before but uh uh agent based and there are no difference.

Basically this second architecture we can get rid of the of the aggregator and this becomes sort of a peer-to-peer or um how was it called gossip based verifiable federated learning where all the agents push stuff and fetch stuff from the from the medium so from the transcript And uh they verify the actions that they perform each other and at the end they can agree on uh the aggregated model uh with a checkpoint with a sort of uh let's say with a sort of consensus and then if the final model and the final I mean if the expected model with the expected characteristic has been reached that's the final model or they can repeat and go on with another iteration exactly how the federated learning uh framework works. So that is to to wrap up how in future uh clients can use uh the federated learning framework to uh to improve themselves for common task. Okay. So this concludes my talks this my my talk. So thank you for your attention.

If you want connect with me, you find my link in Telegram and whatever on that QR code. And that's the preprint of a paper that we recently made where I've taken most of the content uh that I showed you. So, thank you for your attention.

Automatic transcript — names and jargon may be misspelled.