Self-sovereign genomics: turning biology cypherpunk | Péter Szilágyi
Ethereum Cypherpunk Congress·Sun, Aug 9, 2026, 12:00 AM
Speaker
Analysing genomics is currently in it's early phase, dominated by a handful of institutions. As sequencing costs are dropping, more and more people gain access to their genome, but soon realise they have nothing to do with it. Whilst much of the research itself is public, there is no place where that materialises in a usable / reusable way. This results in a sort of soft-gatekeeping, where proper interpretation and guidance is restricted to those with the knowledge or the financials. We're trying to out-build this position, creating a sort-of ""open source movement"" of genomics, puzzle pieces that would allow different ecosystem participants to contribute their small slice of knowledge and build up a global, shared platform for everyone to tap into and extend. Speaker Péter Szilágyi https://x.com/peter_szilagyi Follow Web3Privacy Now: 🌐 https://new.web3privacy.info 𝕏 https://x.com/web3privacy 🦋 https://bsky.app/profile/web3privacy.info 📸 https://www.instagram.com/web3privacy_now/ 🎤 Neocypherpunk Summit: https://s26ber.web3privacy.info/ Subscribe for more talks on privacy, cryptography, digital rights, decentralized infrastructure, and the future of the open web.
Transcript
Uh hey so I guess uh listening to a talk on biology at a web three conference is very very unorthodox but uh just rest assured I also a tech person. My background is definitely cryptography and stuff like that. So um I'll try to keep it a minimum. However, that said, there is one picture I wanted to share. Uh you probably seen it in middle school or high school.
It's essentially cell division. And the really neat part about this is that every single it happens with almost every organism on this planet. It happens with almost every one of us that whenever a cell divides, whenever we grow, whenever we reproduce, essentially we start out with one single cell and that single cell duplicates its internal in its all its internal mechanisms, its internal machinery and then just splits into two. And one cycle later, it splits into four, eight, and you have an exponential growth. And uh essentially a human baby when it's born it it's made up of about three trillion cells but it all started out with one single cell.
And that has some very very profound implications specifically that even though it started out as that one single cell that cell had to contain the entire blueprint for that human body. And uh the implication is that if we can somehow access that blueprint, if we can somehow understand decipher that blueprint, we have or hope that we will have the capacity to solve all genetic all inheritable diseases. Now if we want to do that, it's kind of we need two steps. One of them is we need to somehow get that information out of the cell biologically and the second is we actually have to interpret it somehow. And um the first part seems very very hard and it's actually solved.
So the first human genome was sequenced in 20 sorry in 2003 it took 13 years about 2.7 billion dollar investment. However, ever since then sequencing costs really plummeted. So currently if you want to go to uh if you go to lab you would like to have your genome sequenced it costs approximately $400. Their actual material cost is about $100.
So from a practical perspective, sequencing is solved. Now, of course, we're at a privacy event and uh most people here probably are not really comfortable with just going to any random lab to have their genome sequenced. For those people, the solution is not ideal. You can actually do it at home. It requires weeks of preparation and learning about it and about 2,000, sorry, $20,000 worth of equipment.
Now we can say that okay $20,000 it means it's not solved but u compared to 2.7 billion it's a fraction and if you are willing to wait five more years the trajectory is going down so even though it's not necessarily solved it will be solved so it's on a good trajectory however so from I I mentioned I mentioned there that this is the hard part because this is the physical part Once you actually can take your genome out and sequence it then the only thing you need to do which is very very simple you know just decode it just interpret it that's the easy part and the easy part looks like that ever since DNA was discovered uh we have an exponential growth in research uh publications these are the the papers that contain the word DNA published on PubMed so essentially we have an explosion of data available. Now, one problem is I uh selected one specific paper. It's it was written by one of my friends. Uh you don't necessarily have to read it or understand the title.
But what I wanted to get at is that even though we have an explosion of information being published about DNA, it is so immensely deep and so so very very requires very expert knowledge that there are barely a handful of people in the world who can actually do it. So my friend has been studying I genomics for like 15 years. Um there aren't many people in the world who can invest 15 years of their time to understand it. And this leads us into a very very interesting problem. Um if you would like to have your genome analyzed, if you would like to interpret it, if you would like to know what your genome, what kind of secrets it holds, currently the solution is well you just go to your geneticist and ask them okay you ask your geneticist.
Now guess how many papers out of the 50,000 that was published yes last year they actually read probably five at best. So there's no single person to ask. And uh what they will do is they will just go to a company and just buy a genetic software that can analyze your genome and just get give you back results. But still you still have the same problem. You Let's say you find a small company.
A small company that does genomics will have 20 doctors on board. A big company will have a 100. And if you go to one of the big ones like Novartis, Rosh, they will maybe have thousands of geneticists on board. But the moment the only 15 people in the world can understand one specific paper, what's the probability that they have a person on board that can understand every single thing that was ever published by genomics and on genomics? And my guess would be that it's kind of zero.
You have no chance of finding anybody in any company in the world that can meaningfully analyze your genome from A to Z everything. And uh this is an interesting problem and the thing is we've been here before. Specifically in the world of software not so long ago if you bought an IBM mainframe then all your software kind of came from Micros IBM and or if you bought an Apple Macintosh then Apple gave you the all the software. If you bought a Windows PC then Microsoft wrote you all the software. And the thinking in the world was that well they know what software you need.
They they shipped all the word processing. They shipped all the spreadsheet. They shipped everything you needed to do your work. And retrospectively that sounds very very funny when you think about it when you when you think from the lens of the open source movement of where we are today that at some point everything was developed by one company and we thought it's enough. Obviously it was not enough and is the same thing with genomics.
Currently everything is developed by one lab or a few labs and the general thinking is that this is enough and obviously it is not enough and the question is can we do somehow better can we somehow scale this and my proposition is that what if we could recreate the the world of open source this open source movement but for genomics now if you want to understand it however I think we need to kind kind of unpack it's very hard to say what exactly open source means and for me at least I tried to identify kind of three pillars which are important for me. One of them is uh sovereignty. So up until the point when you every software that you ran was uh was on somebody else's machine on some university lab nobody really experimented. So the whole experimentation with software with open computing started when you could actually buy your hardware and you could actually yolo. You could just tear it apart.
You can do whatever you want with it. That's when people started to actually tinker with it. Now that still was not enough. Um when kind of the normal pieces appeared Windows, Apple etc. what they realized is that well if they want other people to write software on it they have to open up they have to create a platform where all of a sudden it's uh it's not just about uh you can run some program but it's about you can extend it.
They give you pieces of puzzle on which you can build. That was the second very very important thing. And of course the third one, people are uh vain creatures. If we make something cool, we want to share it. So the whole open source movement kind of became possible when we had the internet, when we had the the Bitbucket, the whole sharing site, etc.
Everything when we could simply send our stuff to somebody else to run, when somebody else could actually build on our stuff, that's when the open source movement really started to happen. And uh the question is what stops this thing from happening for genomics because you know open source works so why doesn't it just work for genomics and the answer is in my opinion is that we have a few wrinkles which we didn't quite get right. One of them is irre mistake irreversibility. So with software if you mess up you start over. If you lose your credit card you get a new one issued.
If you lose your private key, hopefully you can roll it. With genomics, if you lose your DNA, it's gone. I mean, it's nobody's getting back. It's once it's leaked out, it's out. So, it's very, very irreversible.
The second is the target audience. With open source software, we built the platforms for ourselves. Essentially, nerds building for nerds. Now, in the case of genomics, we have a split. Nerds can build a platform except we are completely clueless about genomics.
and uh my friend can dissect an eyeball but he's absolutely clueless or he doesn't actually care about building a platform and you have this this huge gap which you somehow need to cross. And lastly, um this is not necessarily specific to genomics, but um even though the internet allows us to transfer data, to transfer our results, to transfer to share in that we can just send each other our results, but sharing it and being able to say that oh that's cool and being able to react it. It's a social thing and that was kind of for example in the current world, GitHub is one of those places where we kind of associate with open source movements. We say that well open source yeah it existed before GitHub but you know kind of after GitHub it took a completely different turn and then the question that I was trying to answer here is that what can we do how can we recreate the open source movement but for uh for genomics how can we somehow um iron out those wrinkles and make this whole thing possible and for that uh just attacking each of these features independently. So the reason open source in my opinion the open source movement became possible is because we became able to tinker with our data to tinker with our machines.
Now if you want to tinker with our genome then we have to make mistakes affordable. We have to make them cheap. We have to somehow make certain mistakes impossible so that if we just screw around we don't have irreversible consequences. And uh for that one option is that yes you can definitely run stuff on your laptop run stuff on your phone that's fine if you are good enough. If you are not good enough, then my proposition is similar to the to hardware wallet in the crypto world where you have a specialized device that has certain security guarantees that has certain um protections in place which makes it extremely hard to misuse, extremely hard to leak your data out because once you have one of those device that you can control very very closely, then it doesn't matter what you do on it.
it it the device should protect you from messing up. It's and that's kind of the important thing to have sovereignty and uh tinkering capability. Now that's let's say you can do whatever you want with your data. The second problem is extensibility. Um we are software developers.
I can write a library. I can there are a few genomics libraries out there but they are not built for biologists. So biologists as I said they care about biological concept they care about an eyeball they care about different cells they don't care about compression algorithms they don't care about cryptography they don't care about any of these things so what is very very important is we want to be able to define APIs that are extremely high level that are targeted for biologists but at the same time be able to um to offer the same protective mechanisms again The idea is that it's very very easy to just create something that's very powerful, but if it's you're going to lose your data with it, then nobody's going to use it. Whereas if you can can achieve the same um uh security restrictions while providing very high level APIs, then you have a chance of convincing biologists to go wild. And um that's for example just one of the um one of the tests that I ran on on my genome.
It's a cute test whether it can tell you whether you are a you like cilantro or you hate cilantro. It's um there's there's actually a genetic marker which can tell you u that part. And um here the only thing that's kind of important that I want to get across is that if you can if you have a device or your laptop that can do um that can do a lot of um that can do all the data processing in advance and all the data indexing in advance then you can have very very powerful applications. And if you can actually sandbox the whole thing and prove that nobody can ever run off with the data, then you can play or you can actually have an AI go wild and play, which is an important thing in the current world. Lastly, uh sharability is uh one thing that's currently missing is there's no place to be able to share genomic apps.
So if I make something, it's mine. If some my biologist friend makes something, it's his. there's basically no way to cross it over. If you can define a platform which has secure execution, sandbox execution, secure data storage, everything and you can create a little platform on top kind of maybe similar to the Ethereum platform or maybe similar to the app store. You get two very very interesting things.
One of them is you get the simplicity of anybody being able to publish these apps. So they can just go wild and publish it. And you also have the extensibility, the composability that open source gave us where one genomic app can actually reuse the results of another genomic app. And this will potentially lead into um these biologists not having to reinvent the world, not having to constantly re uh recreate uh everything. And uh essentially these three steps, these three paradigms are what uh what the the whole Dark Bio project is building out.
It's for the most part uh functional. You can uh actually find me to to have a demo if you want to. And uh for more information um uh we also have a white paper published. Um so that was a very very quick round up. Thank you very much.
And uh find me on outside for more infos.
Automatic transcript — names and jargon may be misspelled.