The Blockchain Native Database by Sebastian Lorenz
Devcon·Tue, Dec 9, 2025, 12:00 AM
Speaker(s): Sebastian Lorenz Event: Worlds Fair Stage Follow us: / efdevcon , / ethereum , https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Devconnect is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon ARG was held in Buenos Aires, Argentina on Nov 17 - Nov 22, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/
Transcript
Good morning everyone. Today we are going to look at what happens when you [snorts] treat blockchain data as a first class domain and build a system around it from the ground up without compromise and from first principles fully leveraging the latest generation of data processing building blocks. Let's start with the reality of the domain we are operating in. First of all, blockchains generally do not expose a clean read primitive. Instead, we are often forced to reconstruct state from raw appendon logs where the vast majority of the data is irrelevant to our request.
We are doing this under constant uncertainty. Blocks can be replaced, transactions can disappear, and chain history can reorganize. But users expect low latency and realtime application interfaces with perfect correctness. Two things that are fundamentally at odds here. This combination makes blockchain data different from what traditional data systems were engineered for.
When we look at a diff a difficult problem as engineers the initial reaction is to evaluate existing wellestablished technologies but most of the traditional data processing technologies were built and designed long before the blockchain. They generally treat consistency corrections like reorgs as an afterthought or a failure recovery scenario. So in absence of a natural fit, we then try to bend these tools to fit our requirements or we end up wiring together a bunch of heavyhanded offthe-shelf components to approximate a solution instead of solving it sustainably. Plumbing does not sustainably solve our problem. Instead, we just end up trading the complexity of the problem for the com increased complexity in our system.
Our problem demands a carefully engineered solution that solves our data needs holistically and sustainably. So, building a database from the ground up used to be a close to impossible feat. It used to require deep pockets, [snorts] a large team of engineers, and years of engineering to pull off. I am happy to report though that that is no longer the case. Our collective understanding of the problem domain and our tools have evolved to the point where building bespoke databases is now encouraged, if not imperative.
The time is finally here for a purpose-built blockchain native database. And this is why we are building AMP. [cheering] AMP is our take on building a database specifically for the unique requirements and the semantics of the blockchain. It provides high throughput batch querying. capabilities married with low latency streaming.
AMP speaks standard SQL for queries, streams, and transformations, making it familiar to data analysts, app developers, and LLMs. AMP builds on top of Apache Data Fusion and Arrow. Data Fusion is a toolkit for implementing custom databases. With it, we can parse, plan, and execute SQL with leading performance at as proven in industry standard benchmarks. Arrow is a data standard suitable for both in-memory data crunching and binary transfer over gRPC.
AMP is where the blockchain data model and the data ecosystem converge. Data fusion does not only benefit us as a foundation to build on top of. It also means that we immediately unlock an entire ecosystem of compatible tools for you to integrate. Arrow and Aeroflight are both welldefined and are supported in a wide range of languages. Additionally, we are building a driver to integrate with the Arrow database connectivity standard ADBC which will enable an even broader connectivity and syncing capabilities.
AMP is designed with developers in mind. It is built as a vertically integrated system. You can either run AMP from a single Rust binary and operate it locally or you can operate it at large scale in a globally distributed deployment. AMP naturally integrates with your local development environment and also the tools that you are using day in and day out. So whether you are using foundry or heartad once you deploy your smart contract to your local node AMP is able to serve you rich data from your smart contracts immediately.
No custom imperative logic with annoying indexing workload deployment cycles. just you, your smart contract, AMP, and a SQL statement. And data declarative languages like SQL tend to really age like fine wine. When you write declarative SQL, you are focusing on the what, not on the how. AMP handles the how when it plans and executes your query.
We do all of the optimization and parallelization for you. You focus on what you want and we focus on how to give it to you in the fastest way possible. Declarative SQL is self-documenting and it's expressive. You know who likes declarative expressive languages? Coding agents.
So to warm us up, here's a very simple SQL streaming query over a raw blockchain data set. reorgs out of the box as part of the streaming semantics of the system. So whenever a new block is mined, it will emit only that new data straight to you. However, AMP is also more than just data. It is compute for computations which are cumbersome to express in SQL.
Data sets can declare and run userprovided functions that may define completely arbitrary logic through these SQL UDFs and can for instance decode on the order of magnitude of millions of smart contract events per second. So so far we have been taking baby steps as far as our SQL queries go. But AMP can also crunch large and complex analytics queries. This query, for instance, extends far beyond the edge of the screen, and it is a fairly involved multi-step query. First, it normalizes token metadata to prepare for joining.
Then, it decodes ERC20 transfers with an early semi- join filter. This filters out non-token transfers before we do the more expensive operation of decoding and then it decodes the filtered locks. Here it extracts the token addresses and prepares for the final join. And then it finally performs a join with all of the token metadata to build the final output. So if you ever need this more than once, we have got AMP data sets for you.
AMP materializes your data transformations into data set tables. Data sets are the domain knowledge of your engineers encoded into declarative SQL expressing the domain model of your application. These tables can then be queried or composed in even higher level abstractions and transformations and reused or composed in other data sets. These data sets can then be versioned and they can be uploaded to a registry where they are continuously streamed and materialized into physical tables. So what does such a registry look like?
Well, we have got an early version of a browserbased query playground online. This playground is backed by such a registry. It is backed by a registry of both raw and userprovided derived data sets. We are opening access at this week's hackathon. So hackathon participants who are participating there can upload their projects data sets.
But be aware of course this is all work in progress. So now we have our data sets. Let's talk about how you can consume them. We provide a reference implementation written in Rust for a durable streaming client on top of our reorgare SQL streaming protocol. The streaming protocol is client authoritative which means that the state is owned by the consumer.
This means in turn that the server does not have to hold on to any state. That is great for scalability and operational efficiency. The stream is also durable which means that it can be recovered and resumed from state checkpoints in case of failure scenarios like networking issues. It comes with transactional guarantees and with exactly once semantics. We're also shipping a TypeScript implementation of the same streaming interface and we'll bring it to other languages in the future.
And because building streaming client clients can actually be quite fun if you have the right abstractions, we've also gone ahead and built a CDC streaming API around the transactional stream that I just showed you. These things are natural to do with AMP and all of its core primitives. AMP is a powerful abstraction for your data and it does not compromise on performance. We achieve subsecond latency at chain head latency not throughput. As far as the extraction and the ingestion throughput goes, we are comfortably keeping up with even the fastest chains on chain head like arbitum and Solana.
This number obviously depends quite a lot on the extraction mechanism. So it depends whether that is a streaming based interface or a polling based interface. AMP can load and decode millions of events per second. We benchmark this with the unis swap swap events and there the decoding throughput at different cardalities peaks at around 4 million events per second. This is a realworld test that includes the fetch, the filtering of all of the full raw blockchain data data set logs and the decoding.
This is the query that we use for this benchmark.
[clears throat]
retrieving all of the Ethereum block hashes since Genesis.
Automatic transcript — names and jargon may be misspelled.