New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Deterministic AI for Prediction Market Resolution | Lucas Martin Calderon - Kash

Ethereum DenverMon, Mar 9, 2026, 12:00 AM

🚀 Get Ready for ETHDenver 2026! 🚀 We're already hard at work preparing for next year's biggest Web3 event! Keep your eyes peeled for more info on ETHDenver 2026—it’s going to be epic! 🌟

Transcript

Hello, Lucas.

Welcome, welcome. Hello, everybody. So, now we have Lucas here. He's going to be talking about deterministic AI for prediction market resolution. Are you ready?

Of course. Yeah.

Okay. Well, stage yours yours.

Amazing. Thank you. So thank you uh well thank you everybody for being here. Uh today I'll be talking about a couple of buzzwords. Obviously I in 15 minutes I barely have enough time to discuss in depth what it means to deterministic AI but you should definitely walk away with you know what are all the necessary elements required for deterministic AI uh to have clear as well.

Hey can I do deterministic AI with commercial models. So all of that jazz actually you would have a much much clearer idea and that's actually the objective of this talk and of course then I will be discussing how it applies to prediction markets and how we want to completely reshape uh how oracles are made for resolution. So to get straight into the into the jargon uh I don't know if you were victims of these markets. I was into one of them. Um won't be able to say which one though.

What do all these markets have in common though? The Zelansky war wears a suit to Trump inauguration or Ukraine agrees uh to Trump mineral deal before April. All of them resolved in a way that people didn't expect even though the rules said very clearly what was going to be you know the the instructions for that resolution. And what is the problem with that? Well, we come back to we use natural language uh for these things.

And when natural language says the word suit may mean different things, right? And there is no perfect solution for this. Uh however, there is something that takes away and removes 99% of all the problems. And that is exactly the paradigm shift of just like in the blockchain in permissionlessness or when everything is you know decentralized you can query whatever service you can query unis swap at any point. Why can't we query the oracle resolutions right at poly market at any point during the open trading uh uh phase right um and that's exactly the paradigm shift that uh that must happen in the field and that's exactly as well what we are developing at uh at cash the company that I lead in in prediction markets and it gets down to if you query the oracle of hey uh whatever service that is right hey cash uh what if um Zelinsky were to use this very exact suit, right?

What you expect from the Oracle is to tell you the exact same thing that will you know the exact same outcome that will resolve upon resolution, right? You don't want the Oracle to say actually that's a suit. If you give an image of a suit and it turns out that is the exact same image at the end and then on resolution the Oracle says oh sorry you know now it's it's not a suit right something very small has changed and it's not a suit. For this you need re uh reprocessibility right you need to be able to replay time and time again the exact same input to get the exact same output and that is determinism right you would hear online a lot of things on chain especially of we've got trustless AI or we've got verifiable AI many of these protocols that are doing I'm sure amazing work they miss something very very important and that is it is very nonsensical to say we've got verifiable AI inference when in fact the inference itself is not deterministic. Right?

And if you want to get deterministic AI inference, you need to get to one objective and one objective only. And that is every time you play with the same seed, you must have bit by bit the same output. Of course, this is very complex. How can you perfectly uh com uh compute every single bit in a in a logic gate that is going to get activated? But there are definitely ways and I'll show you in a bit how we make it happen.

But before I talk into more of the nitty-gritty of you know the the settings the parameters that need to be activated uh let me give you a a very high level architecture of the Oracle system uh that that we are developing. It would be most likely open source so that everybody trusts it uh open for other protocols to use but of course with cache leading the user experience that enables this permissionlessness and this trustlessness when it comes to resolution. Right. So we begin on the left with the AI model. It could be open source, it could be closed source.

Open source is ideal because you can better manage the parameters. You don't have to trust even that the request is going to change in the pathway in the distribution. Uh you don't have to trust as well the response for the provider, let it be open AI, that the parameters, the temperature, all these things are what they're telling you. So obviously uh uh closed source is is not ideal. uh and then you execute all of that within a TE.

Today actually you would be surprised by the level of compute power that SGX has but obviously you cannot execute a larger model on an SGX. you would rather prefer a TDX uh approach and that TDX gives you two things a decap at the station of the integrity of the computation and the actual output of the LLM model right and these two things go into a ZKVM uh you recomputee the entire test station from scratch in the ZKVM it could be risk zero uh they've been doing amazing job it could be succ actually it's a super crowded space so you can find everything uh right and left and uh the and it would output a stark a zikk stark that is about actually that's wrong it's not 200 megabytes it's about 200 kilobytes a thousand times more than a zk snark that uh could be wrapped around through an adapter through gro 16 and then proven on ethereum or any layer 2 based on the EVM uh onchain using the Ethereum attestation service or whatever. So this is probably the nitty-gritty that I was talking about and the single most important slide of of this session because it tells you for once it is quite difficult to find online uh these things or even in academic papers but it tells you in detail what are the specific parameters that we need to have deterministic AI inference. And focusing here on the software uh in the middle uh on the software section there are several uh uh main points but the main ones are four which is the decoding policy which includes you know the temperature the PRNG which is a pseudo random number generator. I'll explain exactly how it fits into the stack and how that serial random uh number is generated within an LLM.

It's quite important to know the basics to know how everything works. Uh how the sample configuration works. You know there is the top K the top P all these technical terms that pretty much mean am I choosing the top 50 tokens uh right after you know the softmax function or am I choosing the top tokens that amount to 90% probability right which may be 20 may maybe 200 right and then the PRNG is the actual algorithm that chooses the next token. So, so yeah, if you want to if you're using if you're not using uh local AI models at home or in your company and you have to submit a request for to conduct deterministic AI inference, most likely you would like to wrap and package all these parameters in a container digest, right? which includes the GPU architecture, Blackwell, H200s, H100s, uh the seed, the driver, the kernel that you're using within the GPU, the decoding policy, and all these elements that I've discussed.

So now let's let's talk about a few misconceptions of the temperature, the PRNG numbers. Many people think, well, if I put the temperature down to zero, it's going to be deterministic the output, right? And well the answer is it doesn't suffice for the temperature to be down to zero and nobody wants the temperature down to zero apart from of course certain commercial cases where you're extracting information from a document or or or things like that. Um but if you definitely want enough fluid intelligence to to be happening in the model you need some sort of uh you know after the sentence of there is not a single path to get to the destination right to get to the solution. It is the exact same thing in in LLMs.

Uh on the right there is a circle you know there is you know let's imagine that circle amounts to the corpus of knowledge of after the training of the LLM. And obviously right in the middle there is a gap right there is uh there there is a gap in knowledge. It could be a time in history that it hasn't been written about. It could be a question you know it could be the answer to a question that doesn't make sense or is uh or or you know it's it's to something that hasn't happened yet. Anyways, there are gaps all around.

So, if you're asking certain questions right very close to that gap layer, right, to that gap horizon uh and you have certain points, certain questions that you're asking, it is very easy for the model with a temperature down to zero to fall into a cycle, right? So, if you're predicting the next word, you're blindly going through a room, right? Uh just touching if you if you can touch anything with your hands, right? and you're going full speed thinking that if your hands are, you know, are not touching anything, you can go even faster, right? So, you might end up falling, you might end up falling through the window, right?

You're going blind in this in in this sense. So, uh so yeah, temperature definitely gets you um yeah uh there's a much higher probability that you end up in a in a cycle. Um and of course there is no creativity. You develop very brittle intelligence uh so to say. So, so yeah about PR andGM sampling the second most important parameters when doing deterministic AI inference is um for for to understand that it's important to see what happens under the hood right at the beginning you have the input make America great you know and then the text must provide and must guess what what the next word is as you know every token is between three four uh letters uh it changes but generally that's the case it first goes through attention layer then it goes through multi-layer perceptron.

Then it goes through attention again, time and time and time again, right? Until you get to the logits, you get to the nonlinear functions of max top k top p. And at that moment when you have an array of let's say 200 different numbers which represent maybe 150 words, you must choose the top 50 or those that amount to 90% likelihood or 10% likelihood. And then from those PRNG chooses the word again, right? uh and even though it was random the way the word again was chosen it is a pseudo random algorithm so were you to run the whole thing again you would get to the word again once more right and that's the point of PRNG when it comes to deterministic inference um and there is to we're a bit tight on time but to to resume as well what are the biggest problems if I don't set up these parameters right is well you get to a point where uh you get to floating floating point non assocusitivity right and everything about AI mostly everything every computation is about matrices it's about multiplicating uh uh different matrices together and depending on how you sum matrices and arrays you can get to a very different output at the very end if you atomically round up uh and you do that sequentially right as opposed to uh summing up everything together right in the case of a b plus c or a plus b plus c in parenthesis right it changes completely It happens throughout the entire life cycle of the LLM from kernel scheduling to variable batching to KV cache memory and KV cache memory is essential for this sort of thing for the attention mechanism.

So let's very briefly in the last three minutes talk about how we are building a whole new Oracle system when it comes to resolution right and even though it boils down to four main points there is almost a year worth of research uh and of course I'm standing on the shoulders of giants in this case uh from people like Andy Hull or uh people uh like uh Carpathy they have worked on very similar methods to create a multi-agentic reasoning system but anyways The fir the fir the the main four points is the single model echo. the fact that even if you divide the same model into tasks to plan, to execute, to deliberate, to observe, to self-critique, as long as it feels within the same context, you're prone to way more uh uh way poorer results, right? So the quality may will be way way less. AI has different weaknesses inherit to the infrastructure as well. Um you know, choosing the next token if you don't have any rack system, it is very difficult to know whether you're going in the in the right direction or not.

um LLMs have a very hard time knowing that they're living on 2026, knowing which model they're running. So obviously these things you must be very very careful when selecting engineering the prompt uh according to these weaknesses. There is overconfidence without selfcorrection again unless they look online or they have a rag system and sensitivity to procedure and framing. And I saw a tweet the other day actually that uh summarizes that quite well. You can ask an LLM who is Tom Cruz's mom and probably would tell you Mary Lee South.

But then if you ask who is the son of Mary Lee South, it would have to look online. It doesn't have the answer. Right? This is to say do not treat the prompts as if the LLM was a database. It has a completely different uh um a completely different infrastructure.

And this is extremely important to learn how this works. And the the the last slides simply means you know what is the performance between executing on a single model versus executing in an LLM council. L&M council with different stages with deliberation and so on. Uh and it turns out that there is a 99% accuracy once you actually add roles personalities persistent memory reputation and anonymized models. So yeah that should be it.

Feel free to ask me later at any point. DM's always open. Thank you very much.

Automatic transcript — names and jargon may be misspelled.