New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Reading Ethereum's Tea Leaves with Xatu data | Devcon SEA

DevconThu, Oct 9, 2025, 12:00 AM

Demonstrate how we collect data from the Ethereum network and how it's used for upgrades, research, and analytics. We'll then run through some examples of how to use the tools and public datasets yourself. Speaker(s): Toni Wahrstätter, Andrew Davis, Sam Calder-Mason, Leo Bautista-Gomez Skill level: Intermediate Track: Core Protocol Keywords: Layer 1, Consensus, Testing, observability Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] hey good morning everyone uh thank you for coming to our Workshop today uh before we start if you have any questions just throw your hand up it'll be pretty relax so yeah uh just look into the workshop overview uh introduction that's us right now uh we're going to go into zatu Genesis um see sort of how it happened why it happened uh and then we'll look into the data sets that have emerged since um we'll then progress into our how we're using the data how others are using you uh using the data uh Leo over here will then present uh he's from the meal Labs team uh he'll present the Big Blocks test that happened last year uh and then we'll have a live tutorial from Tony uh he'll go through how he uses P2 to run his analysis sweet so a little intro on who we are I'm Sam that's Andrew he's hiding uh we've both been doing devops on the panda Ops Team for the last few years is uh and we both have a deep appreciation for things like ethereum and observability as for our team eth Pand drops uh we started out in 2021 uh being thrown straight into the deep end with the merge uh we're embedded in the F and do devops for the protocol uh and we post blogs semi-frequently on our website we try to keep them really high signal um they're the best way to keep up to date with us scan the QR code and we'll jump you straight there our team has a pretty wide range of projects cooking at all times you've probably come across a couple of them uh for example if you've ever run a node you've probably checkpoint synced from an mpoint that was running checkpoint z um you may have also see parry and Barnabas wh been C in the cev calls uh yeah there's a lot of stuff going on so let's move on to the workshop to set the scene uh it's late 2022 and the merge has just happened um we've switched from proof of work to proof of stake um but in doing so our consensus mechanism has become a lot more sensitivity uh sensitive to time uh suddenly the when of things happening has become a lot more important uh and this D this data at a global and network level doesn't really exist um it's easy to check that a block was seen but it's much harder to see when that block was actually seen in Sydney compared to Berlin for example um and now that we're clear from the merge researchers start hacking away they want to upgrade the beacon chain uh but yeah they need this timing data to validate their ideas and and they start capturing it themselves um varying scales different implementations um it's it's hard to expose it's hard to validate there's a bit of yeah potential errors in there so we started to brainstorm ideas on how to solve our problems um we needed to somehow integrate with existing Beacon node implementations as neither of us were really too Keen to implement a full Beacon node um as devops Engineers would usually just Implement a few Prometheus metrics put the feed up and call it a day um but millisecond level Precision is pretty important um so that rules out Prometheus metrics uh log aggregation was also another option but it's definitely a moving Target uh these things change all the time it's not really something that the client devs really pay too much attention to um and we'd also just have to turn on debug level logging um we'd be throwing the kitchen sink at our log aggregation Pipeline and it would potentially be an unreliable result anyway uh so that rules out logs so we started to look at other options uh turns out that the beacon API has this thing called the event stream you can subscribe to it and when the beacon node sees things or does things it will emit an event so blocks attestations voluntary exits everything it was all there uh the Beautiful thing is that the beacon node implementation all supported this uh endpoint in a standardized fashion so what we landed on was zatu uh we used go grpc uh and we thought that it would re be responsible for just collecting ethereum timing data uh it definitely wasn't going to be trying to store or query that data um but the the plan was to derive events and ship them somewhere else to do this we initially created two modules uh zatu server would create events from other modules oh would collect events from other modules and send them somewhere else and zatu Sentry was our first module uh it would run as a sidecar next to every Beacon node uh subscribe to events and send them off to zatu server we wanted to make sure that all of the events followed the same structure so that it was much easier to add new events into the future and also really importantly uh since this is like a distributed system we wanted to make it clear how much you could trust the data so data coming from a client is not necessarily trusted but if it's running if it's been derived by zaru server something that we control uh maybe you can trust a bit more uh this example event is for one of our zaru Century nodes running on mainnet uh subscribing to a beacon node uh and a new block has just come in uh I've redacted a couple of the fields but yeah that's the general idea that's all great but we still hadn't really solved where to send the data um and it turns out it was a lot of data fortunately for us click house exists uh and it is goated uh it allows us to store events at a scale that we didn't really think was possible possible um in July we hit four commas of attestations in our click house cluster with just 14 terabytes of dis usage which is just ridiculous uh given that those attestations we don't keep the signature because we're getting it past the trusted Point uh but yeah still it's it's crazy right data sets I'll hand over to inury data sets um so we just started with Beacon API events and sort of over the last 2 years we've been uh trying to collect it all all the data as a project name suggests um there's there's some overlap between the data sets um for example the beacon API event stream and the lib P2P data sets sort of collect uh events at the head of the chain uh as they're happening um whereas canonical sort of Beacon event and execution events uh is sort of the finalized data um which can be referenced from other data sets so there's a lot of sort of cross uh sort of queries that happen um the there's also so the module Sam talk about like the Sentry um and the lib PP sentury uh we run lots of them uh we collected lot of the same event from lot of different regions around the world um and that sort of gives us propagation times um timing data uh Regional data sort of what's happening in Europe what's happening in us what's happening in Australia sort of allows you to uh sort of differentiate sort of what's happening around the World um we also collect uh me relay data sort of bids uh proposer payloads um we we have a sort of uh Discovery module that sort of crawls the execution layer uh captures mempool from different regions around the world um and we also capture uh canonical execution sort of data um we like to use open source projects uh a couple here so uh hermies by Pro blab uh is a sort of lib P2P uh Sentry I guess you could call it that you run uh connects directly to the lib PDP sort of Gossip sub Network and you can sort of capture that data uh right off the network so it's it's pretty good um and we also use cryo by Paradigm um it allows us to get the sort of canonical execution data pretty good we just uh sort of backfilled all the way to Genesis basically and made that public a couple weeks ago which is a lot of data um also check out Storm from paradigm talk later today on data sovereignty should be good also check out prob laabs talk uh yesterday on how they do a lot of work around the peer-to-peer network uh it's pretty interesting um yeah so we were collecting uh a lot of data for researchers um we didn't really know what to do with it uh does sort of internal EF researches and a few uh researches outside so we decided to open source it and sort of put in parket files um and make it available to everyone uh you can read more there there's a bit more sort of docs on how to sort of access and read but sort of every uh some some tables will be like per day or per hour depending on how big they are so you can grab for example uh Block events for that date and you'll get the sort of raw data in park file so if you want to do analysis uh it's pretty easy to get started um yeah uh couple months ago we allowed so we have we have a lot of centries around the world capturing uh Beacon API event stream data um it's quite costly because you have to run obviously a beacon node and a touch a Sentry um we sort of allowed the community if they want to to run the Sentry next to there be can know um you can opt into sort of how much what sort of level of uh G data you want to send us whether it's uh uh City uh country continent or none at all uh so it's up to you guys but if you want to sort of contribute we're open to that as well uh it's just good to get data from a wide variety of regions and setups and it just helps research as well and all that data we anonymize so uh you can uh so no one can sort of track what Beacon node you're running anything like that but we also publish that data so you can also check it out yourself as well all right using the data uh this one's pretty interesting we use uh Sig P's block Scout uh block print block print tool um which you can check out there it basically use machine learning to try and uh assign a consensus layer client to a validator index uh we then use that data because we've got all this data we know who's supposed to be a testing we know who's not a testing we can then uh link that to client data and it gives us pretty quick response time to figure out what if there's an issue with a particular consensus client implementation uh because you'll see uh thousands potentially thousands of valid days offline for certain uh client implementation so this is really actually useful uh we've got other sort of dashboards grafana dashboards we use uh for Dev Nets test Nets main net um around block timings and stuff like that that we use internally but yeah there's a lot of monitoring goes on sort of about the network about timing um just gives us another view of the network um there's also a pro blab which I mentioned before so they do some cool work so they have uh uh weekly report that uses Z2 data um it shows block arrival time I think there's also uh the size of the block uh and the distribution there uh check it out it's pretty cool um yeah they're doing great work uh after 4844 uh we had a uh ESP data challenge um which was basically WR a a blog post about the uh the blobs upgrade um there was definitely some cool blog uh post there I just picked out one here which was pretty interesting uh but yeah in the in the future there could also be you know more opportunities for data challenge just keep an eye out um also there's another one from Evan from I'm going to butcher that I don't know if it's Prav or prime v um but he did a pre and post sort of analysis of 4844 which which is pretty cool um using that data uh lastly uh some randoms user data um on E research which is cool so yeah that's basically it thanks for that uh I'll pass it on to Leo to uh talk about his awesome work thank you Andrew um so good morning everyone uh thank you for coming to this session um so um yeah I'm Leo from maps and um by the way uh we have a booth uh on front of classroom e in case you want to drop by and keep um asking questions about um what I'm going to present here um or something else that the work that we do um so yes I want to talk a little bit about this um research that we did with the chatau data uh just to put you in context this uh happened uh last year in 2023 uh we have been working uh with the etherum foundation researchers about um data availability sampling and one of the things that um we wanted to know is how much time does it take to disseminate a huge block in the ethereum network so we were building a simul ulator uh that disseminates um huge blocks uh over um over the network and we managed to you know after some months of work to have some idea about how much time does it take to disseminate uh data of the network uh but then we needed to somehow um validate these results that we were getting with our simulator and so we were um you know scratching our heads on how to do that and uh one of the ideas was okay how about we pay a huge transaction into a huge block in the ethereum network and we you know see how much that propag um how much time does it take to propagate over the network so we kind of artificially injected um and by we I mean the interor foundation uh injected um Big Blocks on the network and then thanks to the chatau um uh you know Sentry nodes and chatau we managed to um see how much time does it take to propagate uh then a couple couple of weeks later I was looking at the data and um you know we were recording the block sizes for the last kind of six months or or something like that and um we noticed that there were plenty of Big Blocks in the network and then at that moment I contacted the the EF researcher so I was working with uh Danny Ryan and danr uh at the time and I say hi you guys are still injecting blocks and they told me no no that was that is over and I said well I am seeing these huge blocks on the network what what is going going on and so this is when we kind of discovered like um that there is quite a significant number of Big Blocks that occur naturally organically in the ethereum network and so this figure that you see here is an example of that so this is 6 months I think from April to August 2023 and we are plotting all the blocks that are over 250 kilobytes uh in in this distribution so please notice that this is a logarithmic scale on the on the y axis right and then you see that we go to up to over 2 megabytes um of size for some blocks now um we also found quite weird to see blocks over 2 megabytes because if you know how the um gas system and the C data works on the blocks um you know that usually you have 30,000 um um gas to pay sorry 30 million and um you have um and and then if you want to store one bite of of data you have to pay 16 gas and so if you make the computation that says that the biggest Block in the network should be only 1.8 megabytes and so um we look at that and but in fact the reality is that if you play for zeros uh then that is for uh gas so yes this is Prett yeah so this is this was last year in 2023 so the question for uh it was whether this is pre then up or after after uh 4844 this was last year so this is before that and um yeah so and then actually in in reality you could actually construct a block that is up to 7 megabytes or so and um so because we had all these blocks uh then we can actually had a lot of fun and try to analyze how these blocks propagating the data and that give us a lot of information about how much we can increase the number of blobs in the future uh for data availability sampling and so on so uh thanks to shatu and the centry notes uh that the uh devops guys deploy over the uh over the planet as Andrew was mentioning you can analyze how things behave in different regions of the world right and so um we took this set of uh Big Blocks and we try to analyze what is the you know the distribution of arrival in different locations so uh one of the things that you um can try to analyze is for example okay how much what is the difference of the blocks arriving between Amsterdam and for example Sydney and you see this more or less like half of a second um for um for this particular case and uh uh you also see that most of the times the block arrives within the 4 seconds uh within 4 seconds which is great otherwise um uh this uh things will not work uh because the slots are 12 seconds so you need this data to arrive quite fast um that was one kind one uh type of analysis that you could do but then uh the other thing that we were actually really more interested is how blocks arrive um according to their size so we put the blocks in different beans in different colors from 0.2 megabytes to 2 megabytes right and then we analyze how much time does it take for those blocks according to their size to propagate over the network and um this is more or less the distribution that you see so you see for example that there is between 200 kiloby to 2 megabytes There is almost like a 2 seconds difference uh right there um for for that difference of of sizing blocks right so that is quite interesting because then that give us information about okay what is going to happen when we activate 4844 right after the dun Fork so uh this was actually really great results because 2 megabytes is actually a lot um uh right now we have uh maximum six blobs uh which is about 700 something kilobytes and so we say okay that's fine if we already have two megabytes blocks that propagate over the network without any problem and we didn't even notice it that means that we can allow for six blobs to uh to run over the network and and and we will be we should be okay okay and then um we deployed tenun and we noticed that yeah that actually works uh and then this kind of data um give us information also about how much further we can go in the future right before we Implement data sharding or Pras or whatever we want to call them um so this is why this kind of data is important for researchers and and this is what um uh why this is important and another thing that I want to show is so we we created this uh shatu dashboard you can access it on the on the on the QR code if you want uh that give us realtime data on um on on data that we get from from shatu so we have a bunch of uh different things that I'm not going to um go into detail but I just wanted to mention uh maybe just a couple of them and then you can go and look at uh at the details later or again if you want you can come at the booth and we can discuss about it um so things you can uh see here um so for example uh this is a difference of how the uh the number of blobs looked at the beginning uh so this was during the first uh three months of um after uh Denon was deployed and this is how it looks today so at the time you had like 56% of the blocks with zero blobs and today you have like 40% so the blue color the blue size here has reduced significantly from the initial uh you know the first months of uh Denon uh up to today so this is the last month uh just just like October um and you have seen that the other colors have increased in size on this pie chart right so this has changed the number of blocks with six six uh blobs keeps more or less constant so in this case I think it was 20% right now we are 21% um so that one is more or less constant but the other ones are increasing in size on this pie chart while the number of blocks with zero blobs is decreasing um Tony has also a very nice post that shows uh how this distribution changes over time not only for the counts of blov but for many other things so I also advise you to check on that one and then you can also uh see in real time the propagation of um of blobs over over the network right so you have all these events that chhat is recording and then uh you can plot all that uh for the last month for example you can plot all the blocks and all the blobs and see what is the uh latency uh of arrival for all the all the blobs in the in the network right again you can access this uh this data on our dashboard and we also have an API in case you want to play with the data as well um another thing that you can do uh is try to look at how much uh blobs are used so that means how much data is actually put inside so when you have a blob and then at the end you have um like 50 kilobytes of zeros then we are assuming that that is not used right and so you can look at those details and and see what is the usage of um of the blobs and uh so this is how it looks from the last month so I think uh they are very fairly good uh use of the of the blobs so that's like almost 90% uh of the space is being used in the blobs uh that was not the case um some months ago uh the the it was something around 60% I think uh some time ago so this is also evolving uh you know as rollups and layer TOS get more familiar with how to use blob data um um they this is changing so again if you were interested uh please look at the dashboard and another question that was uh interesting to analyze was how many M blobs sorry how many M blocks occur after some number of blobs right and this is important because if it takes too much time to propagate a block with six uh with six blobs then what are the chances that the next one is going to be a Miss because the the block proposal was busy the downloading blobs or or whatever and and and so this is something that we we looked at uh this was actually a question that was RA that was raised during the interrup in Kenya and and and that we wanted to analyze and so chatu allows you to do this kind of analysis uh and so this is again one of the pie charts that uh is is shown in our dashboard in real time and you can see that you know most of the blocks that were missed um in the last month uh occur just after a block with zero blobs okay and and again you have the distribution over here um but there is uh not such a strong correlation that you will believe that okay every time you see a block with six blobs then the chances of missing the next one are very very high that's not what you see in this data and again there is a post by by Tony uh on on that so also advise you to look at that maybe he will mention it later um yeah so I think that's uh more or less what I have today and then we can move on to the to to Tony thank you [Applause] I needed okay can you switch to my computer ah perfect yeah hi everyone um so my part will basically be about how I use KATU for my daily analysis and I want to focus especially on Patu which is basically a library um in Python very um if you're familiar with 3. Pi for example this is basically what I tried to achieve here allowing users to have a very high level very easy access to the P to the Cato database and then yeah coming up with some functions abstracting the complexity of KATU although KATU is not very complex but of course you can always make it simpler by by wrapping in within a python library and yeah so let me show you the GitHub so you can find um my GitHub there is a py Patu library and you can install it very easily by simply doing pip install Patu and then you have to run KATU do setup which will copy a configuration file into your home directory there you have to put in your credentials and then you're um ready to go so if you need credentials for KATU I would just say reach out to the to the um e pands team i' um they are very happy to provide you those and yeah so as soon as you have the credentials in your home directory um you can already start by doing something like this um we basically okay now I have to kind of Pardon are the dimming the lights is that possible can we dim the light if that's possible okay let me see how I do you want to hold it okay okay let's see oh it's taped it's taped to the table okay then very nice maybe you want to use that chair okay okay so how would start it's basically importing the KATU database import Patu then I would create a um the object that basically inherits all the functions which is um um Patu do Patu sorry right and then you already see okay I'm now collected to this um endpoint and I will then call this endpoint so in the background it's actually using SQL queries but the python Library should allow you to make that as simple as possible and the first use case that I thought would be uh maybe appropriate for having a quick demo is what if we want to know what are the M slots over time so first thing I would do is I would have to find the right function everything is documented in the GitHub but basically it's kat. getm slots right and of course I don't know which columns are actually behind this data set so there's a data set called um Beacon slots or something similar and in order to know which slot was actually missed the simplest thing to do is we um we query all slots and then we check um Which slot doesn't have an execution payload right because if if there is no execution payload then this means basically there is no um yeah nothing from the El in that block which means either the block was re all the block was missed so what I would do is I would first get the docs from this specific specific function so I would just check okay I need the get Miss I copy it paste it and then I can see all the different columns I can get from this data set that is behind the get Mist slot function for example I'm now interested in actually only one thing I only want the slot number right I don't want to I don't need anything of uh any other column I basically only want to know which slot number was missed so the query would then look like the following I would have um kat.

getet Miss slots first I would specify the slot range that I want to query which is done like this let's say 9 million to 10 million and the second thing is the columns I want and this looks like this and now let's see if this works wonderful so what we get here is basically now a python set with all the slots that have been missed within slot 9 million and 10 million and then I can continue by putting it into a pandas data frame let's do that very quickly oh I haven't imported pandas yet okay um um let's see DF pandas that will be like Miss slots oh I called it DF and then I think I need to specify the column here let's see if that works already looks good okay that's nice now I have um everything in a panas data frame and the next thing I would do is I want to map the slot number to the actual date right this is yeah please um this is all mayet yeah so the yeah I I could mention that um when you query for this for um sorry it's actually here you can also always specify additional parameters I've skipped it now for um Simplicity but basically you could say I only wanted for um a specific test net you could say sepolia heski but um per default it's always mainnet because this is also very important if you query you always have to make sure that you're not mixing different um Nets up with each other right because suddenly you get very strange things I can tell you um where you're like okay is that actually proo is mayet broken but then you find out okay it's actually a test net but thanks for the question okay so the next thing I want want to do is basically add a date column and there you can do the the following so I will use um the um Python's Lambda function that's very handy for use cases like that and then there is aat helper function um let's see how it is called get slot daytime I guess yeah let's lose this one and I would basically throw in the slot number because every slot you know we have 12 seconds so knowing the Genesis time we can exactly tell at which time a slot um appeared let's see if that works wonderful and now we are already quite far so we have the date the slot and the last thing we need to do is then to group by the date um let me quickly check okay let me see get slot okay there is another helper function that would be simple slot today okay actually let me use another helper function because now I got a very precise timing with seconds and everything but in order to visualize this m slots over time we don't need the minutes the hour and everything we basically want to have it grouped by day so there is also slot today right and there I would do the same right that's already a little better and now I want to group that whole thing let's see Group by I want to group by date I want to sum up the slots basically I want to count them actually because otherwise I would just add the slot numbers and then reset index just this is all this is basically all some Panda stuff to get a nice data frame in the end index and in the end we can sort it by date right let's see if that works perfect and then to make it 100% realistic I show you exactly how the workflow works you can already see C gbt is ready I would basically do something like give me uh I like the library plotly a lot plotly Express chart visualizing my Mist slots over de days paste in the the my data that I have and then yeah my um I had quite good experience with chbd so chbd is really good at quickly quickly creating those charts so first let's copy the whole thing so I don't need the sample data that jet gbt gave me but I can see it put it into a DF variable so I would do something like DF is my group DF and then let's see if that works already perfect right so this is so this is um actually very simple example how you get to the m slots I don't know how much time it now took like 5 minutes 10 minutes but it's a very realistic workflow so this is how how it looks like when I do many of my analyzers of course it gets a little more complex um especially if we want to go a little deeper but for the M slots um that's actually already it the next thing we could do is maybe I should ask is there a validator in this room or someone who has a favorite validator validator number one okay let's do Val okay yeah and they are not yet withdrawn okay let's see because one very interesting thing that you will not get anywhere else but you get at exato database is you can check when your blocks are seen in the network so you can check okay my validator proposed the block and I want to know when did this block arrive in Europe when did arrive in the US Australia and so on and this is very interesting especially if you want to check um who is playing timing games um are your blocks late for example if you build your blocks locally you don't use MEF boost then this is uh in particular interest of you because you might want to know okay is my block out there at the network in one second into the slot or 3 seconds into the slot right because if your block is very late then you might want to consider using item F boost or yeah looking for better peers and just to make sure that you will not be be Reed okay I see validator number zero is active let's let's do it with validator number zero so there is a function called sat. get Block event lock event that's the one I will just do the same slot range again which will where do I have it I got it here let's do 9 million to 10 million okay now we have to hope that this validator had a slot in this range otherwise we would just deviate to another one he proposed something he propos he proposed a block in August yeah okay okay then we should find something first I do the get docks again because I'm um this is always handy to see what KATU is actually offering here it's actually do get dos and then I paste in the function name okay what I need now is they event daytime so there are many different noes all around the world different clients everything and I want to know when are blocks seen by those different clients and the block and the time scene is here the event daytime because it means when was this event recorded and because the events are recorded in real time as soon as they are seen we should be able to use this event daytime then let's see what else we might need columns is equal to event daytime we also need the slot I guess and we might need the we also need the proposer ID right we want to know the proposer which validated ID had did the proposer actually have but you can see this is I think not included in this data set so we might need another query for the proposer which tells us when which proposer was supposed to propose in which slot get proposer same range and again that's how it works we go into the docs again quickly check what is there there is a slot and there is a proposed valid dat index and I would then query those two slot plus propos a valid data index and then I should get a pandas data frame that shows me the slot and the proposal index now we can go on and merge those two um data sets so here we had the the block event which looked like a slot and the event daytime okay I should have not run it again okay looks good perfect and then we can do something like pd. merge um proposer we merge um left and we merge on the slot number left on slot and then write on slot looks good right and then we are already getting very close right now we have the slot number the proposer that proposed a block or was supposed to propose a block in that in that slot and the event daytime when we when the block was seen could you the next thing we would need to do is basically group another time because for example if you want to know when it was first seen right this is very interesting because for example if you want to know if someone is playing timing games you want to know the first time the block was seen in the network usually that's in second like one in the slot but if you play timing games then it's like a second three in the slot and what we can do here is basically group the whole thing we Group by slot and I think it works like this regroup by slot and propose a validator index and then we take the minimum of the event dat time oh D okay okay let me quick Qui check let me make sure that I didn't forget anything here yeah it looks good um the maximum would be very different because sometimes a node sees a block let's say 20 seconds after it was actually proposed right for example if the nosis is slightly behind or something so you will find sometimes numbers that are much higher for example let's say the I don't know Nimble SC combination in Australia had some problems right at that time then you might see numbers that are way too high but usually all those notes should agree on yeah some rough number and I'm here interested in the first time um we saw a block because this tells us um with a very high likelihood that this is the time when the block was proposed the block was actually proposed a little um earlier but usually this should be a quite good approximation okay let me check because I just see I I should have done reset index set index um let's do in place true okay no maybe there is no in place let me try it like this oh okay that's fine and now one very interesting thing is um you've all seen on Dune dashboards on my dashboards we always use those valid dat labels right and it's it's super hard because you need those different hristic to create those labels but you can also use Kato in the future you can have you can access it under kat. validators and then. mapping and what this gives you is basically um the valid dat ID public key the deposit address and all the labels um that you see on Dune dashboards or on my dashboards um about all those validators so what we could do now is let's first check um we can let's just merge it to our data set what we have so far which is the same again merge DF we Pro we have the labels we merge on the validator ID so left then we do left on okay the left hand side it was called proposal validator index and on the right right side it was called um validate ID right okay this looks good now we have all that information in our data frame and what we can do now is we can already start looking for um what is our Target validator doing validator with the ID um zero oh okay we want validator ID is equal to zero oh there was no block let's try one okay validate ID one had a block in that time so for example we can see um the time when it was seen um all the other information that we need and finally what we can do is there is another helper function for you that is telling you at which time in the slot the block was seen so seconds in slot and then we do DF do apply so I apply a function over the whole row and then Lambda X and then I have to do to helper helpers and then I will do get time and Slot okay get the time and slot slot and what this function wants is basically the slot number and the actual time it was seen in the slot so the slot number we got it in the slot and then um the event dat time okay actually I forgot one important thing here just see that I should have converted this daytime into a actual Unix timestamp but this is very easy we do it like event daytime is equal to to event daytime and then you can do something like um panas has a function that makes sure that your um field is actually um a date format so daytime we can say UTC is true because the all the times you find in our UTC UTC is true right and then we can do something like S Type which means we convert from daytime to um integer um integer and this now gives us milliseconds we want seconds which means we do something like um yeah 10 to the power of um six this should be good enough let's quickly check if our data set looks like expected the event daytime yeah we now have uh the datetime in Unix timestamp and the final thing we do is exactly this here we continue with X event daytime right um let me quickly check I didn't forget anything looks good okay this is everyone who is familiar with pandas knows why I do certain things for example apply one because we now put in the whole row and not only one cell and let's see if that works okay this is kind of expect that let's see where the bug is oh it's access one another ply yeah right and what it what this does is basically we tell him the slot we tell him the exact time time we uh recorded the event and using that information you can then tell at which second in the slot was the event actually seen and now we have all the information we need for example what we can do is let's actually do something like group by we Group by label and then we do something like seconds and slot we take the let's do the median and we reset index okay let me quickly check I think median without the brackets no Group by label seconds in slot okay there's a typo huh oh I still called it ah I called it event daytime right no I didn't do the seconds and Slot yet right this one ah the S oh thank you thank you so this should work perfect and then we can do median right and then we get the median time when every value dator slot is seen you can see we've got a lot of yeah looks like some solo stakers but then we can also check like okay we all know that P2P ke there are some entities playing timing games right so what we could do is we do something like this label equals let's do K so for K the median time when the SL when the block is seen is 2.

93 and then we can do the same name let's see for example okay no block in that time ah wait then you can see okay that's kind of the honest Behavior if you push your blocks out without playing timing games your blocks should be seen within the first second of a slot whereas if you play timing games you're at two point 2.9 seconds and yeah this is exactly how I did for example analyzis on timing games if you want U more information you can check out my e research um profile I did many posts recently with uh especially pushing for a blob increase in pcture analyzing big blops analyzing reorgs and yeah if you have any more questions or interested in in your own validator into another Pool just let me know and we can look it up thank you [Applause] other questions s thanks for coming perfect [Applause]

Automatic transcript — names and jargon may be misspelled.