NimbleCo AI · Open Science

Beyond the Paper

The scientific paper was designed in the 1880s for a world of print journals, physical mail, and limited space. We now live in a world of unlimited storage, executable code, and live data streams. This is an argument for decoupling what the paper bundles — and a map of what already exists.

The paper bundles four things that should be separable

The scientific paper isn't a single thing — it's a container that welds together four distinct layers of scientific work. When these layers can't move independently, every failure cascades.

The Paper Claim + Evidence + Method + Narrative
locked in a single static PDF
Claimsindependently verifiable nodes
Evidencelive, linked, queryable data
Methodsexecutable, version-controlled
Narrativean optional view, not the unit

When you can't update the evidence without rewriting the whole paper, can't reuse the method without reimplementing it from prose, can't challenge one claim without attacking the whole publication, and can't see when evidence has shifted under a claim written five years ago — you get a system where the narrative (the least scientific part) becomes the unit of credit and citation. The container is the problem, not the people.


Fourteen structural critiques of the form

Each of these is a necessary consequence of the paper format — not a failure of individual scientists, journals, or peer review. They are built into the form. You cannot fix them by asking people to be more careful, because the form demands them.

01

Inevitability framing

The Introduction constructs a backstory that makes your study look like the inevitable next step. The reader never sees the branching paths, abandoned directions, or experiments that went nowhere. Science gets rewritten as a clean forward march.

Fix: Claim-level nodes show the actual path — every attempt, every dead end.
02

Selective framing of the problem

You choose which prior work matters, which gap is "the" gap, why your question is the one that needed asking. The framing is a rhetorical move, not an neutral survey — but it presents itself as the latter.

Fix: Knowledge graphs let you see all connections, not just the ones the author chose to draw.
03

Forced novelty

Every paper must claim novelty. Replications, confirmations, incremental improvements all get dressed up as "novel contributions" to survive review. The literature saturates with novelty claims where most content isn't novel — and genuinely valuable non-novel work is pushed into a shadow economy.

Fix: Novelty isn't the price of admission. Replications update existing claim nodes. You get credit for being informative, not new.
04

Inflation of figures & supplemental materials

Papers have ballooned. Supplementals are often longer than the paper itself, rarely read, unsearchable. The paper becomes a packing crate rather than a communication.

Fix: There's no "supplemental" in a live system. Data is the data. Link to it, render it interactively, explore at whatever depth you want.
05

P-hacking

The paper shows you the one analysis that worked, not the fifteen that didn't. The garden of forking paths is invisible. You see the result, never the search that produced it.

Fix: Executable, version-controlled methods. Every variation tried is in the commit history. P-hacking becomes visible because the forks are preserved.
06

Rich-get-richer citations

The Matthew Effect: papers cited early get cited more, because people cite what's visible and safe. Citation count becomes a popularity metric decoupled from evidentiary quality.

Fix: The metric is quality and independence of supporting evidence, not citation count. Circular citation chains become visible and discountable.
07

Self-citation

You inflate your own apparent impact by citing yourself. The citation system has no mechanism to distinguish self-citation from independent endorsement.

Fix: Network structure makes self-citation visible and computable. "Effective independence" of evidence can be calculated.
08

Speed-vs-rigor tradeoff

Priority — who published first — gets you credit, tenure, grants. So there's systematic pressure to publish before you're confident. Once published, findings are anchored and hard to walk back.

Fix: Findings are provisional with explicit confidence levels that update. First observation gets credit (timestamped), but quality determines status. No single publication moment.
09

Positive results bias

Journals prefer positive findings, so the literature is systematically skewed. The distribution of published results doesn't match the distribution of actual results. The published record is literally wrong about the prevalence of effects.

Fix: Negative results update existing claim nodes. The system shows what was looked for and not found — the information you need to calibrate beliefs.
10

Narrative structure (IMRaD)

Science isn't a story. It's messy, iterative, often boring. IMRaD forces it into: problem → heroic intervention → results → meaning. You must perform significance even when your contribution is modest.

Fix: No template. Contribute what you found. Compose a narrative from nodes for communication — but it's a view, not the thing itself.
11

IMRaD makes you a poser

The format makes you pass yourself off as radical when you're just doing careful work. You need to follow the formula even if your science doesn't fit the formula. Replications, tool papers, null results, exploratory analyses get shoehorned or dropped.

Fix: Each contribution type is first-class: replication node, null result node, dataset node, method node. No shoehorning.
12

"Is this a paper?" before "should the public know?"

The first filter is publishability, not scientific value. If it won't make a paper, it doesn't get communicated — regardless of whether the public or other researchers should know about it. The form is a gatekeeper for what counts as knowledge.

Fix: The question flips: "Is this something someone should know?" If yes, publish it — as a claim, data update, method, null result.
13

Experiments that don't fit get dropped

If an experiment doesn't serve the narrative, it's cut. Those results are lost to the record. The paper format actively destroys findings that don't fit its shape.

Fix: Every experiment is a node. Nothing gets dropped. The narrative is composed from a subset; the rest persists independently.
14

1880s form for a digital age

IMRaD was standardized in the 1880s, widespread by the 1940s–50s. It was designed for print journals, physical mail, and limited space. Why would the optimal form for a world of unlimited storage, instant distribution, executable code, and live data be the same as the optimal form for typewriters and postal mail?

Fix: This is the reductio that anchors the whole argument. The form is the problem.

The IMRaD timeline

A format designed for a world that no longer exists.

1880s

IMRaD standardized

The Introduction–Methods–Results–Discussion structure is formalized in the natural sciences, driven by the constraints of print journals and postal communication.

1940s–50s

Universal adoption

IMRaD becomes the default structure for scientific publishing worldwide. Typewriters, physical typesetting, and limited page budgets define what a "paper" can be.

1990s

Digital arrives — the form doesn't change

Journals move online, but the PDF remains the unit. Supplemental materials balloon as digital storage removes page limits — but the paper's narrative structure stays fixed.

2010s

Open access & preprints

arXiv, bioRxiv, Sciety, OpenAIRE emerge. But they make the paper open — they don't question the paper itself. Access widens; the form stays.

2020s

Executable & live — but still papers

eLife ERA, Stencila, Code Ocean make papers executable. Nanopublications break claims into RDF triples. But each fix addresses one layer; none stitch them together.

Now

The decoupling

We have all the pieces. What's missing is the infrastructure where a live dashboard's output becomes a claim-node in a knowledge graph, with an executable method and a confidence score that updates with replication.

Evidence Overlay

When you read a paper, you don't check every citation. You can't. A citation is a social gesture that signals "this claim is backed" — but the backing might not be good enough, and you'd never know without tracking down and evaluating every cited source for every claim.

The evidence overlay makes the backing visible. Every sentence (or clause) is annotated with an evidence score derived from the actual quality of its support — not "does it have a citation" but "does the cited source actually support this claim, and how strong is that evidence." Toggle between reading mode and evidence mode to see the hidden structure.

Live Demo · Excerpt from a (simulated) Introduction

Strong Moderate Weak Unsupported Unverified

Sleep is increasingly recognized as critical for memory consolidation[1]. During slow-wave sleep, hippocampal replay events reactivate neural sequences associated with prior waking experience[2], and disrupting this replay impairs spatial memory in rodents[3]. Interestingly, recent evidence suggests that sleep deprivation may also affect emotional regulation through amygdala–prefrontal connectivity[4]. However, the mechanisms by which sleep loss directly alters decision-making circuits remain poorly understood[5]. Here we show that 24 hours of total sleep deprivation in humans significantly reduces prefrontal–amygdala connectivity and increases impulsive choice in a temporal discounting task[6]. These findings suggest that sleep loss produces a distinctive neural signature in decision circuits that may underlie real-world risk-taking behavior[7].

Click any highlighted claim to see its evidence trail.
The point: In reading mode, this looks like a normal Introduction. Every sentence has a citation. It reads as authoritative. But switch to evidence mode and you discover that one claim has strong replicated backing, another rests on a single study that hasn't been independently replicated, one is unsupported by its cited source, and another cites something we couldn't verify at all. The citation gives the same visual signal regardless of evidentiary strength. The overlay breaks that illusion. — Citation-resolution is real, but not live: the demo's references were run through Matilde's verifier (Crossref, OpenAlex, Retraction Watch) on 2026-08-05, and each claim shows that dated result recorded in the page — it is not re-checked when you load this. Click any claim to see it. Claim extraction and evidence-support scoring are still a mockup. The pipeline is buildable with existing tools.

The technical pipeline:

1. Claim extraction

Parse scientific text into atomic assertions. Each assertion becomes a queryable unit with its own evidence status.

2. Citation resolution

Map each assertion to its cited source via DOI, URL, or metadata. Verify the source exists and check for retractions.

3. Evidence evaluation

Check whether the source actually supports the claim. Assess evidence strength: replicated → strong, single study → moderate, indirect → weak, no support → none, unverifiable → unknown.

4. Rendering

Color-code or annotate each sentence/clause. Reader toggles between clean text and evidence view. Click any claim for full provenance.

What exists already: Matilde's citation verification tools (Crossref, OpenAlex, Retraction Watch, Unpaywall) already do step 2 for individual citations — they verified every reference in the demo above. Steps 1 and 3 (claim extraction, evidence-support grading) are not yet automated; the demo's extraction and grades are authored. The birdsong dashboard already demonstrates live-evidence display. The overlay is the same principle — making evidence status visible at the point of reading — applied to scientific prose.

The Landscape of existing alternatives

We're not inventing from scratch. Each piece of the post-paper future already exists in isolation. What's missing is the stitching — an infrastructure where these approaches connect. Here's what's out there, what each fixes, and what each leaves untouched.

Approach What it fixes What it keeps Status
Nanopublications
Smallest unit of scientific assertion as RDF triples with provenance. Machine-readable, citable, FAIR.
Claims layer Provenance
Breaks claims into atomic unitsMachine-readable No evidence layerNo narrativeNo UX
Real implementations (KGHub, WikiCite) since ~2010. Adoption limited by RDF complexity barrier.
Executable Research Articles
eLife + Stencila. Papers with live code blocks, programmatically-generated figures, dynamic values. Reader can modify and re-run.
Methods layer Reproducibility
Executable methodsInteractive figures Still a paperIMRaD structureNo claim decoupling
Live on eLife, GigaByte + Code Ocean. Most mature "executable paper" implementation.
Publish-Review-Curate (Sciety)
Preprints published immediately, community peer review and curation happen publicly and continuously afterward.
Review layer Speed-vs-rigor
Continuous reviewNo gatekeeping delay Still a preprintNo claim decouplingNo evidence overlay
36,000 evaluated preprints, 76,000 evaluations from 27 community groups.
Overlay Journals
Peer review layered onto already-public preprints (arXiv, bioRxiv). No separate publication artifact.
Access Cost
Open accessLow cost Same paper artifactSame narrative structure
Open Journal of Astrophysics, Quantum, Episciences. No author/reader fees.
Scholarly Knowledge Graphs
OpenAIRE Graph (250M+ works), OpenAlex (209M works), SciLake. Map relationships between papers, citations, authors, institutions.
Relationships Assessment reform
Maps citation networksOpen citation index Above the claim levelNo evidence scoringPaper is still the node
Active infrastructure. OpenAIRE designed for research assessment reform (CoARA).
Registered Reports
Method and analysis plan peer-reviewed and accepted before data collection. Eliminates p-hacking and publication bias for the registered part.
P-hacking Positive results bias
Eliminates p-hackingReduces publication bias Still produces a paperOnly for pre-registered analyses
Growing adoption across journals (COS registry). Still a minority of published work.
Post-Document Science (AKE)
Proposed "Autonomous Knowledge Engine" combining formal claim verification, MoE routing, and information-theoretic filtering. Shifts epistemic control from narrative to computational verification.
Claims Verification
Computational verificationKnowledge graph integration Theoretical proposalNo working implementation
Published 2025 (MDPI). Conceptual — closest to our decoupled model in the literature.
Micropublications
Extends nanopublications with structured evidence and attribution. The smallest unit of assertion plus its supporting evidence.
Claims + Evidence
Evidence-claim linkingIncentivizes data publication RDF barrierNo narrative layerNo UX
Published concept (Oxford Academic, 2018). Limited adoption outside biocuration.
The pattern: Every existing alternative fixes one layer but keeps the others. ERAs make the paper executable but keep the paper. PRC fixes review but keeps the paper. Overlay journals fix access but keep the paper. Nanopublications fix the claim level but lack UX, evidence, and narrative. Knowledge graphs fix relationships but stay above the claim. Nobody has stitched the layers together. — That gap is where our proposal lives.

The Proposal: decoupled, stitched, live

Don't abolish the paper — demote it. The paper becomes one possible view of an underlying evidence system, useful for communication and synthesis, but no longer the unit of scientific knowledge. The unit becomes the verified claim with its evidence trail. Everything else is infrastructure around that.

Four layers, four lifecycles

📌 Claims
First-class, independently-verifiable, confidence-scored nodes. Each claim has a persistent ID, a provenance trail, and a status that updates with new evidence. A novel claim enters with low confidence and no incoming edges. Each replication adds edges. Confidence accrues. The novel claim isn't privileged — it's on probation.
📊 Evidence
Live, linked, queryable data. Not supplemental files — the actual data, rendered interactively, explorable at any depth. Evidence links to claims; when data changes, claim confidence updates. The birdsong dashboard is a working prototype of this: live data, live analysis, live display.
⚙️ Methods
Executable, version-controlled pipelines. The method IS the method — not a prose description of it. Full commit history shows every variation tried. P-hacking becomes visible because the forks are preserved. Anyone can re-run the analysis on the same data, or on new data.
📖 Narrative
An optional composition layer that sits ON TOP of the other three — not instead of them. You can still write a story, make an argument, synthesize findings. But the narrative is a view, not the thing itself. The nodes persist independently. Nothing gets dropped because it "didn't fit the story."

How this addresses each critique

Every critique from Tab 1 maps to a structural fix in the decoupled model.

Inevitability framing →

The knowledge graph shows all paths, not just the constructed narrative. Abandoned directions persist as nodes with "explored, no finding" status.

Forced novelty →

Replications and null results are first-class contributions. Credit accrues for moving the field's confidence, not for novelty.

P-hacking →

Version-controlled methods preserve the full search history. The garden of forking paths is visible.

Positive results bias →

Negative results update existing nodes. The system shows what was looked for and not found — the information needed to calibrate beliefs.

Citation gaming →

Effective independence of evidence is computable. Self-citation and circular chains are visible and discountable.

IMRaD / narrative →

No template. Each contribution type is first-class. Narrative is a view you compose, not a cage you fit into.


A new model of peer review

The current peer review system has two jobs welded together: gatekeeping (deciding what gets published) and evaluation (assessing whether claims are supported). The decoupled model eliminates gatekeeping — everything gets published — and reinvents evaluation as a continuous, claim-level, democratized process.

Review at the word level

A reader enters review mode on any narrative or claim. They click any individual statement — a sentence, a clause, a single word choice — and write a critique. The critique attaches to that specific claim node, not to the "paper" (which doesn't exist). Every claim carries its full critique history, visible to any reader.

✍️ Critique at any granularity

Disagree with how data is described? Critique the specific word. Think a causal claim is actually correlational? Challenge that clause. Think a methodology section glosses over a confound? Flag that sentence. No more writing a 3-page review of an entire paper — you critique the exact thing you object to.

🔗 Critiques are themselves nodes

Each critique is a first-class node in the knowledge graph — with its own evidence trail, confidence score, and critique history. A critique can itself be critiqued. The graph tracks the full dialectic: claim → critique → response → counter-critique.

Democratized critique: expertise is in the argument, not the credential

You don't need a PhD to have a good critique. A field biologist who has spent 20 years observing kākā in the wild may spot a categorization error that a lab-based ornithologist with a PhD missed. A statistician may catch a sampling bias that the domain expert overlooked. A citizen who experienced a drug side effect may have data that contradicts a clinical trial's conclusion.

The principle: The system evaluates the expertise contained within the critique — not the established expertise of the person making it. A well-evidenced critique from a citizen carries more weight than an unevidenced assertion from a professor. Authority is earned per-critique through the quality of the argument and evidence, not granted wholesale by degree or title.

How the AI agent evaluates critiques

When a critique is submitted, the AI agent doesn't just accept or reject it — it evaluates the logic and evidence of the critique itself:

🧠
Logical evaluation: Does the critique identify a genuine flaw? Is the reasoning sound? The agent checks the logical structure — e.g., does the critique correctly identify a correlation-as-causation error, or is it itself making a logical error?
📋
Evidence evaluation: If the critique provides evidence (its own data, alternative analyses, cited studies), the agent runs the same verification pipeline on the critique's evidence as it does on the original claim. Does the critique's evidence actually support its assertion?
⚖️
Confidence adjustment: The agent adjusts the original claim's confidence based on the weighted critiques. A strong critique with solid evidence lowers the claim's confidence. A weak critique that doesn't hold up doesn't. The adjustment is transparent — the reader can see exactly which critiques moved the needle and why.
📊
Integration into the graph: Valid critiques become part of the claim's permanent record. When the AI agent uses that claim in synthesis or analysis, it accounts for the critiques — weighting the claim's contribution by its post-critique confidence, not its original assertion.

What this replaces

Old model

2–3 anonymous reviewers with PhDs, chosen by an editor, review the whole paper before publication. Months of delay. Binary accept/reject. Reviewer expertise assumed relevant but unaudited. Critiques invisible to future readers. No mechanism for post-publication correction except a separate paper.

New model

Anyone can critique any claim at any time. No delay — published immediately. No binary verdict — confidence adjusts continuously. Expertise evaluated per-critique. All critiques visible permanently. The AI agent handles verification at scale; humans contribute judgment and domain knowledge.

This doesn't eliminate expertise — it redistributes it. The PhD-trained immunologist still matters. But so does the nurse who noticed a pattern across 200 patients, the statistician who spots the sampling flaw, and the citizen who experienced the side effect. The system's job is to evaluate whether each contribution is well-evidenced and logically sound — not to check whether the contributor has the right degree. Expertise is demonstrated through the quality of the argument, not assumed from the title. — This is the structural answer to "who gets to participate in science?"

What it actually looks like: the platform

A web platform where scientific knowledge lives as a living system, not a static archive. Here's the concrete logistics.

The user experience

A researcher — or a citizen, a journalist, a funder — lands on the platform. They don't read papers. They explore claims, each backed by live evidence, each with a confidence score that reflects the quality and independence of its support. They can:

🔍 Query the knowledge graph

"What do we know about X, and how confident are we?" The graph returns claims with their evidence trails, not a list of papers to read and evaluate yourself.

📊 Explore live data

Click through to the actual datasets — interactive visualizations, raw downloads, executable analysis pipelines. Not supplemental PDFs.

📝 Submit a finding

Upload data, register a claim, link evidence. The system handles DOI minting, metadata extraction, and evidence-graph linking automatically.

🔄 Replicate or challenge

Submit a replication or a null result. It links to the original claim and updates its confidence score. First-class contribution, not a "failed paper."

👁 Toggle the evidence overlay

Reading a narrative? Switch to evidence mode and see which sentences are strongly backed, which are weak, which are unsupported. (This is Tab 2's demo, applied to real text.)

✍️ Enter review mode

Click any individual statement and write a critique. The critique attaches to that specific claim node — not a whole paper, but the exact sentence or clause. Anyone can critique; the AI agent evaluates the logic and evidence of each critique. No PhD required to participate.

🤖 Ask an AI agent

Natural-language interaction: "Find all claims about sleep deprivation and decision-making that have been independently replicated." The agent queries the graph, fetches the evidence, and synthesizes — with full provenance for every claim it makes.

The AI agent's role

An AI agent is built into the platform — not as a replacement for human judgment, but as infrastructure for scale. The thing that makes the paper format impossible to replace manually is that checking every citation, evaluating every evidence trail, and maintaining every claim's status is more work than any human can do. The agent does:

📥
Ingest & structure: When a researcher uploads data or submits a claim, the agent extracts atomic assertions, resolves citations against Crossref/OpenAlex, checks for retractions, and links the claim into the knowledge graph automatically.
🔎
Verify & score: The agent evaluates whether cited sources actually support each claim, assesses evidence strength (replicated → strong, single study → moderate, indirect → weak, no support → none), and assigns a confidence score with a transparent evidence trail.
📡
Monitor & update: The agent watches for new replications, retractions, and corrections that affect existing claims. When a retraction occurs, the claim's confidence drops and all dependent claims are flagged. The system stays live — it doesn't wait for someone to notice and write a correction.
💬
Answer & analyze: Users query in natural language. The agent synthesizes answers from the claim graph, fetching and analyzing data on demand — running analysis pipelines, generating visualizations, comparing evidence across studies. Every answer carries full provenance: which claims, which evidence, which confidence level.

The agent doesn't replace peer review. It handles the mechanical verification work that humans can't do at scale — checking 200 citations against their sources, flagging circular citation chains, detecting when a retraction propagates. Human judgment sits on top: researchers evaluate claims, contribute replications, and curate the graph.

Access & sustainability

Free tier

All public claims, evidence, and data are free to read, query, and explore. Anyone — citizen, researcher, journalist, student — can browse the knowledge graph, view evidence trails, and read narratives. This is the public-good layer. The data is open; the knowledge is open.

Token-based AI interaction

Querying the AI agent (synthesis, custom analysis, evidence-gap reports) uses a token system. Users get a free monthly allowance. Beyond that, they bring their own API key (OpenAI, Anthropic, open models) or purchase tokens. The platform is model-agnostic — you use whatever AI backend you prefer.

Institutional subscriptions

Labs, universities, and funding agencies subscribe for: private claim graphs (pre-publication), team workspaces, custom analysis pipelines, and monitoring alerts. This funds the free tier. Subscription replaces journal APCs — and costs less.

Funder/grant integration

Funders can require that grant-supported findings be submitted as claim nodes with live evidence (instead of papers). They get real-time dashboards of what their money has produced. Research assessment shifts from counting papers to counting verified claims.

The economic argument: The current system costs ~$10,000–$25,000 per published article (APCs + subscriptions + reviewer time + editorial overhead). For that price, you get a static PDF that decays. The platform provides a live, verified, self-updating knowledge artifact for less — because the mechanical work is automated and the infrastructure is shared. The money moves from maintaining a print-age pipeline to maintaining a live knowledge system.

Technical stack (what we'd build with)

Knowledge graph

Claim nodes in a graph database (Neo4j / RDF triplestore). Each node has: persistent ID (DOI-like), assertion text, evidence links, confidence score, provenance, version history. Built on OpenAIRE / OpenAlex infrastructure.

Evidence layer

Data stored in object storage (S3 / Zenodo). Interactive viewers rendered server-side (the same pattern as the birdsong dashboard). Executable methods in Git repos with CI runners.

AI agent

Model-agnostic backend. Citation verification via Crossref, OpenAlex, Retraction Watch, Unpaywall (already built — these are Matilde's existing tools). Analysis pipelines via containerized code execution. Natural language interface via API.

Frontend

Web platform (React/Next.js or similar). Evidence overlay as a text-annotation layer. Knowledge graph visualization (D3.js / Cytoscape). Live data dashboards (the pattern we already use). Mobile-responsive.


We're already building this

The pieces aren't theoretical. They exist, they work, and they're in production:

A live evidence system: real data, real analysis, live display. No paywall, no gatekeeper. The "publication" IS the running system. nimblecoorg.github.io/open-science/experiments/nz-birdsong ↗

Citation verification tools

Claim-checking against Crossref, OpenAlex, Retraction Watch, Unpaywall. The evidence-overlay pipeline is these tools wired into a text renderer. Already operational — used to verify every citation in this dashboard.

A structured claim graph: each scope-line is a claim node with its testability status. Not biased by citation count — biased by evidence. nimblecoorg.github.io/open-science/experiments/consciousness ↗

Community-generated data with valence analysis — a working example of non-expert knowledge contribution, structured and queryable. nimblecoorg.github.io/open-science/experiments/peptide-reporting ↗

The question: Do we build the infrastructure that makes the stitching possible for others — or do we keep demonstrating it case by case until the pattern becomes undeniable? — This dashboard is itself the argument for building it.

Repository of alternative approaches

A living reference list of existing attempts at post-paper scientific knowledge. This will grow.

Nanopublications

nanopub.net · RDF · Since ~2010

The smallest unit of scientific assertion as machine-readable, FAIR digital objects using RDF semantic web technology. Each nanopublication contains an assertion, provenance, and publication info. Real implementations via KGHub and WikiCite. Limited adoption due to the barrier of formalizing claims in RDF.

claimsRDFsemantic webFAIR

Micropublications

Oxford Academic · Database · 2018

Extends nanopublications with structured evidence and attribution. Incentivizes community curation and places unpublished data into the public domain. Similar to nanopublications but with richer evidence modeling. Limited adoption outside biocuration.

claimsevidencebiocuration

Executable Research Articles (ERA)

eLife + Stencila · Since 2021

Published papers with live code blocks, programmatically-generated interactive figures, and dynamically computed values. Authors use R Markdown / Jupyter with Stencila Hub. Reader can modify code and re-run directly in the article. GigaByte + River Valley have pushed further with Code Ocean integration.

executablemethodsreproducibilityeLife

Stencila

stencila.io · Open source

An office suite for reproducible research. Author interactive, data-driven publications in visual interfaces similar to conventional office suites, built from the ground up for reproducibility. Underpins eLife's ERA format.

executableauthoringreproducibility

Sciety — Publish, Review, Curate

sciety.org · Since 2020

Preprints published immediately, community peer review and curation happen publicly and continuously. 36,000 evaluated preprints, 76,000 evaluations from 27 community peer evaluation initiatives through 18 organizational partnerships. Aggregates preprints from multiple sources for discovery and evaluation.

reviewpreprintscommunitycontinuous

Overlay Journals

Open Journal of Astrophysics, Quantum, Episciences · Since ~2013

Peer review layered onto already-public preprints (arXiv, HAL, Zenodo). No separate publication artifact. The Open Journal of Astrophysics is fully funded by Maynooth Academic Publishing, no author or reader fees. Cost-efficient compared to traditional subscription publishing.

accesscostarXivopen access

OpenAIRE Graph

graph.openaire.eu · EU infrastructure

A scholarly knowledge graph of 250M+ works with open citation links, designed specifically to reform research assessment away from journal-impact-factor metrics. Embodying POSI (Principles of Open Scholarly Infrastructures) and CoARA commitments. Active infrastructure for open science in Europe.

knowledge graphcitationsassessment reforminfrastructure

OpenAlex

openalex.org · Since 2022 · 209M works

A fully-open scientific knowledge graph launched to replace the discontinued Microsoft Academic Graph. Metadata for 209M works, 13M disambiguated authors, 124K venues, 109K institutions, 65K Wikidata concepts. Free, open, with API. Not yet at the claim level — maps papers, not claims within papers.

knowledge graphmetadataopeninfrastructure

Registered Reports

Center for Open Science · Multiple journals

Method and analysis plan peer-reviewed and accepted before data collection. Eliminates p-hacking and publication bias for the registered analysis. Growing adoption across journals via the COS registry. Still produces a paper at the end, and only covers pre-registered analyses.

p-hackingpublication biaspre-registration

Post-Document Science: Autonomous Knowledge Engine

MDPI · Information · 2025

Proposes Object-Oriented Scientific Information (OOSI) and an Autonomous Knowledge Engine (AKE) combining domain-specialized MoE routing, formal verification of claims, and information-theoretic filtering. Shifts epistemic control from narrative evaluation to computational verification. Conceptual — no working implementation yet. Closest to our decoupled model in the literature.

claimsverificationcomputationaltheoretical

Publish-Review-Curate Model

Plan S / Coalition S · 2023

A model where publishing (preprint) happens first, review and curation follow. Sends a clear signal that the community has evaluated and validated the article. Endorsed by Plan S as part of the transition to open access. Decouples the timeline of publication from evaluation.

reviewpreprintsPlan S

The Case for Reform of Scientific Publishing

International Science Council · 2021

Comprehensive roadmap for reform: improving peer review, ensuring open access, shifting from "publish or perish" to valuing diverse contributions, and prioritizing global dissemination of knowledge as a public good. A high-level policy argument that the system is broken — but doesn't propose decoupling the paper form itself.

policyreformISC

Existing infrastructure we can build on

These aren't post-paper alternatives, but they're the plumbing our decoupled model would connect to:

Crossref

DOI registration & metadata

Persistent identifiers and metadata for 150M+ scholarly works. The backbone of citation resolution.

Retraction Watch

Retraction database

Tracks retractions across the scientific literature. Integratable via Crossref for automated retraction checking.

Unpaywall

Open access resolver

Finds legal open-access copies of papers by DOI. Essential for evidence access without paywalls.

Wikidata

Linked open data

General-purpose knowledge graph in RDF. 100M+ items. Potential substrate for claim-level nodes.