Data is Creative Energy
New Foundation·Thu, Jul 18, 2024, 01:19 PM · 16:51
Organic and synthetic intelligence create net-new dynamic knowledge and advance our societies, economies and the state of the art, forming a new exponential asset class. Anna Kazlauskas (Vana), Jason Yhao (Story Protocol), Saneel Sreeni (Ritual), Erika Mann (Newfoundation) Covering insights on: • User owned data and user-owned foundation models • IP on-chain for stronger security guarantees and to provide attribution, provenance, & traceability of creations
Transcript
Okay, guys, thank you so much for joining us during the lunchtime. And welcome to our guest, I'm Sofiane. And yeah, let's start with a little round of intros. Maybe also we can mix it with a question, which is, what are you building and do you think that data is the new NFT? Hey, everyone, I'm Jason.
I'm one of the co-found Hey, everyone. I'm Jason. I'm one of the co-founders of Story. And a little bit about me, I spent two years as a product manager at DeepMind before co-founding Story. And in terms of what Story's trying to do, which is very related to your question, Story is building the world's IP blockchain.
And one of the issues that we really ran into, I ran into at DeepMind, was we were training on the entire web, right? And part of this involves, of course, the data of users, you know, the sort of things that you produce on Reddit or on Yelp. But also, really important was that we were training on data from creators. And one of the things that Story focuses on is ushering in a future where creators and their IP can be monetized in the age of AI. And we don't believe that you can really cease and desist every single AI model that's using your creative style or your voice.
I mean, sort of, especially if you're Drake or you're someone big like JK Rowling, that IP is already out there. So the question is, instead of playing whack-a-mole with internet, how can you work with it? And as a creator, how can you allow AI to leverage your IP and allow communities to extend your IP through AI and so what we focus on a story is on ramping IP to on chain and making it programmable in the same way that circle has on ramp fiat on chain and made it programmable and DeFi has done it for the entire financial ecosystem the question that we're asking and trying to answer through our technology is can you on ramp creativity and allow creators set the terms almost like a robots.txt for IP for how AI, how other programs, how other applications can leverage and extend that IP and compensate creators properly. And the way that we're doing that is through a standard that leverages both the 721 standard as well as a 6551 token-bound account so that every single piece of IP on story, not only does it have a real legal license backing it, it also has an NFT to represent the media file, and then this token-bound account, which represents the logic or the rights.
Because IP, at the end of the day, is not just a static media file. It's media plus rights. It's media and money, and that's what the smart contract account allows us to do. Hey, everyone. I'm Sunil.
I'm a founding member of Ritual. Maybe to borrow Jason's analogy, if they're making IP programmable, we're making models programmable in sort of a blockchain-native way. So Ritual is building an L1. We're building our own sovereign chain that essentially allows you to access models natively anywhere on the chain. And then we build in a lot of semantics for the people who are running these models around how do you get rewarded for running your model, how can applications be routed to use you, how can you fold in cryptographic schema around integrity and proofs.
And for us, we sit very, very close to the application developer. We expose models. You can write smart contracts generally over models, so we expose models to them. So the data is a little bit more upstream, I would say, in the supply chain, but it's still very, very important, right? Where there's very general purpose models, and then there's models that are very, very well suited for specific domain tasks, which require a lot of proprietary data.
And so when you think about maybe data being the new NFT, I do think that we will see a massive, massive shift towards proprietary data, because what we've learned is that no matter how good a model is, we can only teach it to do things, but we can't teach it to remember from training data. At the end of the day, you still need to give it contextual information from maybe a vector, from retrieval augmented generation, like vector databases, etc. Or you need to try and, like, you know, a very specific architecture over a very specific corpus of data, at which point you lose your general compatibility or your general purpose compatibility. So to that end, yes. I wouldn't say it's the new NFT, but I do think that there will be knowledge stores that are fairly non-fungible with each other and they will be very valuable.
Now, what value is described to that is a much harder question. Thank you. So my name is Art. I'm one of the co-founders of Vana. We are a network designed specifically for people to earn the value of their personal information or their personal data by putting it on chain and being able to trace how it's used in AI models.
One of the projects that was built on our early testnet was the Reddit Data DAO. I don't know if you've been following some of the press around that lately. net was the reddit data dow i don't know if you've been following some of the press around that lately um in a partnership with us and aura the reddit data dow has now uh launched their first on-chain model uh trained on user-owned data um i think that's relevant to the question of is data the new nft i think one of the things that makes this a really difficult uh question is that there are certain properties of data that we don't yet know and certain economies that data will create that we don't yet know of. So to think about the framework of an NFT as it applies to data, I think would probably limit what that would be. And so that's why we've decided that we're going to create the L1 that allows for these properties to kind of permeate as the data is transformed in different ways.
So for instance, the Reddit Datadow could have decided to sell their dataset. Reddit actually reached out to the Reddit Datadow and said, hey, we don't like what you're doing. And the Reddit Datadow responded and said, well, change the way you're dealing with our data. So data can be used as a form of social activism. And data can be used to directly create a model.
So now the community owns their own model. So I think that is data the new NFT? I think some aspects of it are, but I think that we don't yet quite know what we can do with our data in the markets that it's creating. Awesome. Great answers.
Thank you, guys. Yeah, basically the idea of the NFT as the universal way to represent something, like the unit of accounting of creativity, let's say. And which leads to the next question, which is also a hard question, sorry about that, which is, you know, when you look at data as like an abstraction of human creativity, how do you measure that? So in the real world, we have market mechanisms like the music industry the fashion industry the art sector where you have like different protocols in place to like define the value of of that creative output uh you know the stock market being another one uh yeah do you see it as like a market mechanism or is it more like using ai uh as a way to kind of measure the value of that creative output from humans and also maybe from machines? I'll start on that one.
I actually come from a data sales background, so I could probably speak to this from a Web2 perspective. So how your data is currently exchanged or traded right now is that platforms take your data, they use third-party providers to scrape or to try now is that platforms take your data, they use third-party providers to scrape or to try to access that data, and they create a market value based on how difficult it is to undertake that task. It's a very opaque market, it's a very closed market and it doesn't let you actually accrue the value of your own data that's being scraped or being committed to these projects without your knowledge. So I think there's going to be a lot of market discovery and what the value of data is worth once our projects kind of come to light. But that shouldn't – I mean, every new marketplace has some form of discovery based on the market.
I think one of the critical things, one of the critical proofs of the VANA ecosystem is called proof of contribution. And it's the idea that when you add a piece of data to a data set, the value of your data relative to someone else's should be the value that it contributes to the overall data set. And we're doing some research with the University of Toronto on model influence functions as one example of proof of contribution, which basically takes your piece of data, runs it through a function created by a model and says, how much does this piece of data, runs it through a function created by a model, and says, how much does this piece of data differ from something that's already there? So if somebody had uploaded it beforehand, obviously the value would be quite low. So I think that there's a lot of market discovery to be done in this space, but that's the reason why we're doing what we're doing to break these monopolized markets and to break these kind of closed markets so that we can find value in our individual data.
Yeah, market discovery is a very, very interesting thing. I'll speak to this maybe from the model perspective and someone who's spent a little bit more time there. I spoke earlier about contextual data and how that's important both for the training of models that are good for specific domains and for injecting contextual data. And so one thing I think quite a lot about is for specific proprietary data, that is actually very, very valuable. And now with things like retrieval augmented generation, vector DBs essentially, you have a way to monetize contextual data without necessarily forking it over.
So if I have contextual information around me as a person or Jason or a specific domain, anyone with access with the model and a compatible encoding scheme can go and pay me to access that corpus of data to augment the inputs, right? And the only way they will ever discover the full extent of my data is if they somehow manage to find every single prompt that would traverse the entire vector DB. And so when I think about market mechanism discovery, for me at least, I think that's where I think we'll start seeing very interesting things where we might very well see marketplaces around contextual information and vector DBs in the coming years. And what value is ascribed to that basically will come down to what value that additional data, you know, holds in terms of augmenting, you know, inference. Now, beyond that into training, et cetera, it's much, much harder to tell.
And I think these two are far more qualified to speak on that. It's a really interesting thought with the vector DBs. You know, ultimately, like your question around will AI price data or will markets price data, it sort of boils down to the familiar debate of the 20th century in economics, which is the socialist calculation debate. As technology gets better, can we have a central planner that creates pricing? And I'm pretty skeptical of that.
I think that in general, markets have proven pretty robust. And I do think what you mentioned, Art, about the ability for models to give, the ability for us to figure out how much a piece of data is contributing to a model's output, I think that will be a useful signal into a market structure. I find it very hard to imagine that that tool or that information itself is the price, because price is so endogenous. So that signal will become part of the market. And so I think that these technical innovations will make contributions.
But what we're doing is creating markets. And then also, these sort of technical contributions are signals into the markets. As far as how Story thinks about it, from a protocol level, we actually don't focus on pricing. What we do is we allow creators to set the terms for what they think their IP should be worth. Now, of course, there's obviously a lot of friction there because most creators tend to overvalue their IP and that's good because they should be confident in what they're creating.
And so we think that there will be a second layer of applications like marketplaces that help creators or maybe use AI to set an initial starting point for creators. And then there can be mechanisms like auctions that help us price IP accurately. But all of that is an improvement over how IP, at least from a creative perspective, is priced today, which is a bunch of agents having behind the door closed deals in a market of around 100 people, right? And maybe like five to 10 big players. So ultimately, even though I don't think pricing will be solved immediately, I think the combination of technical signals and market forces and allowing creators to start and attest to their own IP's price as a beginning point, I think that will get us most of the way there.
Awesome. and attest to their own IP's price as a beginning point, I think that will get us most of the way there. Awesome. Thank you so much. Let's maybe wrap it with one final question for all of you.
How do you see this basically roadmap from where you are starting and what's the kind of north star? Where do you want this to achieve? Yeah, for us, IP is a very complex asset, right? And with money, the main value is store value, medium of exchange. And I think that comes together pretty well in Ethereum and in Bitcoin.
The way we think about IP is that we have sort of a funnel or roadmap in terms of how IP is onboarded and how we make it useful. So what happens is the first step is, can we actually just bring IP on-chain? Can we create markets around this? Because IP is actually a multi-trillion dollar asset class. It's larger than, at least as of now, all of crypto combined.
And it's just one of the most easily understandable asset classes. You don't have to explain to someone why Snoopy is valuable or Mickey Mouse is valuable. Everyone owns IP. If valuable or Mickey Mouse is valuable. Everyone owns IP.
If you have an Instagram account, you own IP, right? But at the same time, it's this extremely opaque, extremely liquid market. So the first step is just bringing IP on chain, tokenizing it, allow it to be programmable. And for every, let's say, 1,000 IPs we bring on the story, maybe 100 of them will be actually licensed, right? So we've built out infrastructure to automate licensing with a single click.
Maybe 100 of them get licensed and people are creating comic books around your IP or adding specific color or clothing to your character, whatever it may be. And then out of those 100, maybe one to 10 of them will receive royalties. So there's a sort of funnel where we're bringing IP on chain, allowing people to program it, and then people will license it or remix it, and then some of that will generate revenue. That's the process of bringing the entire creative ecosystem on chain. But I think what we've learned with crypto is that oftentimes liquidity is really important and trading and exchange are really important in order to bootstrap that network.
So how we think about the ecosystem is very much tokenization and then licensing and then royalty and monetization. Yeah, I think for Ritual, right, our whole point is that we can create this sort of like shared layer where anyone can go host or you know utilize hosted models build applications over them and have all the um right rewards and incentives in place for that right um and i think for us very simple i'll start with the north star is that i want every single model every single powerful model in the world in like 20 years to be hosted on ritual i think it'll be the best place to come and host a model. It'll be the easiest for developers to come and build on you. But the way we think about where we are right now is that we've already released a product that's fully cold-frozen that was our phase one for Infernet that allows people to access models on-chain. And basically, based on some of these semantics, you can do payments on-chain.
If you have an internet connection, you can get to Ethereum or any EVM-compatible infrastructure. You'll be able to go and access models that are hosted on this network of nodes. That's already done well. We've done a few million inference requests over that network, on-chain and off-chain. And the next stage really is the ritual chain, where we basically take a lot of this and encode it as native infrastructure so you don't have to call out to an oracle this is like accessible to any developer on the chain right um and for us i think like the very much the bet is that you know ai obviously is a fast-growing technology of all time right models are sort of the crux of this is where they all value actually flows from is like who's using and paying for models there is a massive gap right now in the market where we have very very performant open source models um they're very, very difficult to monetize until you build an application around them versus also very performant closed source models, but they're highly concentrated.
It's like maybe three or four players at best that people go and use them, right? So we can build the rails where it's one, highly easy to go and build over all these other models and even proxy to these big models that people want, which is largely the best way as a smart contract. You can think of any blockchain as one big open API, one big open backend where everyone can go and interact. So if we can do that and then we can fix the problem around does it make sense to go and run an open source model because how am I going to drive flow to it? And then three, solve all of the trust issues around privacy integrity.
Then we basically would have just created the best layer for all models to exist on, which is why we say we're the execution layer for AI. For us, that's really the end vision. We're working very, very closely with partners both to augment our infrastructure. We work with Story around model IP along with a lot of applications that do everything from edge inference to consumer to DeFi ML. And the goal is really to start with a subset of the market that makes a lot of sense for people that are really aware of this crypto AI intersection and want to use them.
Once you can prove that out, I think anything's possible. I was telling Jason this earlier today, but our worst case is suddenly that Ritual becomes the way to pay for model access cross-border, which is still probably within the next five years a trillion-dollar market. So I'm going to be short in terms of what our North Star is, and I'll ask the audience a question. Who here has used a large language model? Who here has used a closed-source large language model?
And who here owns a piece of that closed-source large language model? And that is what we're going to change. We want there to be a collectively-owned foundation and who here owns a piece of that closed source large language model. And that is what we're going to change. We want there to be a collectively owned foundation model because our data has gone into that.
And we want to envisage a future in which the technology that we use is technology that we own. And that's simply it. That's the North Star. Wow. Amazing, guys.
Thank you so much. Great to chat with you. And, yeah, thank you for attending. Thank you so much. Great to chat with you.
And yeah, thank you for attending. Thank you.
Automatic transcript — names and jargon may be misspelled.