New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Mapping and Funding Our Dependencies in Open Knowledge

ETHBerlinThu, Jun 19, 2025, 02:14 PM · 10:20

Open knowledge has diverse forms — including open-source software, scientific research, and collaborative data commons — that are all built upon intricate webs of dependencies. As vividly demonstrated in cascading OSS failures caused by overlooked maintenance of dependencies (e.g., the npm left-pad incident), sustainable open knowledge production requires ongoing care and investment in its foundational contributions.

Transcript

Okay. Sorry. There we go. Okay. Can we go back one?

There we go. Okay. All right. So, let's go. Hi, I'm Warren.

Yeah, I'm with DRIPS Network, and yeah, I'm going to talk about how knowledge graphs and continuous dependency funding mechanisms can help with the alignment of incentives in research and open source for more reliable work. So, it can be easy to forget, but the knowledge work really is a game of deep interdependence. We build and we reason on top of layers that are created by others' efforts, open source code, shared data, scientific discoveries. The Stewart Brand quote, information wants to be free, it's true in the sense that the distribution costs of data approach zero, but producing reliable knowledge does require investment in social coordination and validation work. When those foundations are neglected and layers fail, the impact can be systemic.

For example, when Babel, Node, and thousands of other packages broke in 2016 due to the retraction of the 11-line dependency left pad, or last year when the XE backdoor nearly compromised SSH globally. Like this KCD classic pointed out, it's absurd that so much common infrastructure depends on the maintenance of obscure software by unpaid volunteers. And just as overlooked software dependencies can trigger cascading failures, flaws in foundational reports, data sets, and protocols can propagate costly errors in research. This is worsened in science by especially laggy corrective loops. Journals often take years to retract discredited publications, and retraction notices can take roughly three years to appear on PubMed.

Since retroactively correcting papers isn't explicitly rewarded, those zombie citations readily embed into the literature, like Andrew Wakefield's fraudulent report linking the MMR vaccine to autism that ended up getting cited over 1,200 times after it was retracted. The NIH has estimated that every retraction costs around 400K in wasted funding, and that doesn't account for all the costs of downstream research. This is all part of a broader alignment issue in the scientific community that's rooted in a publish or perish dynamics, where the participants have to optimize for primarily the novelty and number of publications they're turning out. And this is selected for some strange patterns, like only 3% of psychology journals explicitly inviting replication studies and clinical trials that yield null results, having roughly one-third the odds of publication of clinical trials that can reject the null. Or that the peer review and citation patterns disproportionately favor novel results, which creates disincentives for publishing confirmatory or null results or for performing resource-intensive validation work like triplicate studies.

Researchers also often withhold their data and their code out of fear of scrutiny or getting scooped in highly competitive systems that severely punish errors. My intent, by the way, is not to scapegoat researchers at all. It's only to point out some enabling factors that have led to a broad reproducibility crisis that's especially pronounced in cancer biology, psychology, and neuroimaging. And I really know that to be true. I started out my career in a couple of those fields.

And while some funding agencies like the NSF and NIH have been pushing open science mandates, enforcement's uneven and coordination failures persist, both because of institutional inertia and, I think, because of an issue Friedrich Hayek framed as the local knowledge problem, where effective resource allocation depends on context-specific insights that are dispersed across networks and are outside the reach of centralized bodies. Luckily, we have some nice decentralized funding mechanisms today like RetroPGF and deep funding and DRIPS that help address that limitation of – by drawing on the domain expertise of maintainers and researchers throughout networks, which can help calibrate incentives at the ground level. In the case of DRIPS, the protocol taps distributed local knowledge through what we call continuous dependency funding, where funding is streamed incrementally and allocated by programmable routers called DRIP lists. DRIP lists define how streamed funds should be split proportionally across up to 200 recipients per list, which can be GitHub repos, Ethereum or Filecoin addresses, ENS names, or even other DRIP lists. These lists support recursive funding flows, where when a recipient claims funds, they're encouraged to reallocate some of that funding to their own dependencies through their own DRIP lists.

Over time, those re-splitting decisions generate this valuable public dataset of maintainer-source dependency judgments that reflect their actual perceived impact in their ecosystems. The protocol is extensible to fund any off-chain identifier that has publicly accessible metadata that our oracle can fetch and use to verify address ownership. That's how GitHub repos are able to participate on the network. And we're soon expanding it to allow researchers to participate through their HuggingFace repos and their ORCID profiles. ORCID, which is the increasingly standard open researcher identifier.

But one issue that the protocol doesn't inherently address is the cost of eliciting crucial contexts that experts have on which dependencies are most impactful to them. Maintainers and researchers intimately understand their dependency hierarchies, but manually translating that into DRIP lists comprehensively takes due diligence and time that's often crowded out. And that challenge has motivated us to start building algorithmic support for specifying DRIP lists. We believe that algorithms should accelerate and assist experts' dependency judgments, not replace them. In the prototype, we initially map a repo's local dependency graph through static analysis, and then we score the external dependencies through personalized page rank plus an LLM's interpretation of user-supplied contexts, like about dynamic dependencies or the business logic importance of certain components.

The resulting DRIP list provides an initial funding split that maintainers can quickly refine and optionally feed back to the algorithm and hopefully save themselves a lot of time in the process. This approach extends naturally to research using citation graphs like OpenAlex to surface high centrality entities. But at the algorithmic step, we'd like to avoid relying on commonly good-hearted metrics like the H-index, since that might just encourage amplifying existing misalignments. Fortunately, page rank index is a pretty well-motivated approach for scoring researcher impact that tends to dampen the influence of citation farming and strategic self-citation patterns that game the H-index. This, I think, should only be a starting point, though, for a few reasons.

The first is that page rank index can still be gamed through collusion. Another is that citations vary in their functions. Citations that declare intellectual debt should probably be weighted more heavily than ceremonial and strategic citations, and that requires edge-level labeling of citation intents. The third is that some key dependencies, like datasets and software and protocols, are dark matter in citation graphs because they're routinely omitted from bibliographies. So, what might be the shape of a knowledge dependency graph that brings better visibility to the research substrates that need more support?

One is that its nodes would be made up of both explicit dependencies, like the ones declared in bibliographies and manifest files, and the rest that can be conservatively machine-inferred, like inline citations of software, protocols, and datasets. I think another is that its edges should be able to be filtered down to only those representing genuine dependency relationships, which citation intent classification can help with, using, for example, the SySyte dataset. A third, slightly more ambitious property would weight each downstream dependent edge by a reliability score that's inherited from its source node. We have something like this for open source, thanks to public vulnerability datasets and software composition analysis tools. For research, I'd look to draw on the replicability predictive modeling work that the Center for Open Science has been sponsoring, and in the meantime, I'd start with simple heuristics, like preregistration and data transparency flags.

Altogether, these three properties should enable a kind of transitive reliability assessment for research, so that funding can get directed to fragile foundational nodes that are surfaced along paths to high-impact topics. There are some non-trivial challenges, like the methodological heterogeneity across fields and what to do about closed-access papers, but I still believe that this is an especially high-leverage path to align research incentives with greater reliability. So, I'd love to hear from you if you're interested in collaborating.

Automatic transcript — names and jargon may be misspelled.