The bottleneck in research has moved. AI agents now produce hypotheses, algorithms, and candidate solutions in abundance — the cost of a conjecture has collapsed. Validation has not. And part of validation was never physical: knowing how a machine-generated claim was reached, how confident it was, and whether that reasoning can be re-examined. That is decision provenance.
Karl Popper described science as conjectures and refutations. Agentic AI has broken the symmetry of that pair. Conjectures are now cheap and near-infinite; refutations remain physical and institutional — slow, expensive, and paced by the laboratory. The validation deficit is widening, not closing.
Beside the well-known wet-lab bottleneck sits a quieter one: the epistemic-audit bottleneck. A hypothesis a reviewer cannot trace is a hypothesis a reviewer cannot trust. A faithful log of an unexplained machine decision is still an unexplained decision — faithfully logged. What is missing is the provenance of the decision itself.
Leading AI-for-science labs have converged, in 2026, on a concrete set of requirements for research agents. In their own terms, agents:
Each of these requirements describes a record — one that must survive hand-off to a peer reviewer, a funder, or a regulator, and be checkable without access to the system that made it. Decision provenance is the open, verifiable form of that record.
A Causal Seal binds one AI-generated output — a hypothesis, a candidate compound, an algorithm, a proof sketch — to the causal parameters that governed its generation (which model, under which constraints, seeing what context, driven by whom) and the moment of emission. Anyone can verify a seal's integrity with no access to the model, its weights, or any private system.
That is, concretely, the interaction card the field is calling for: portable, model-agnostic, and tamper-evident. It travels with the claim from the agent that proposed it to the human who must decide whether it is worth a month of laboratory time.
The most trusted scientific models — structure predictors are the canonical example — report a calibrated confidence for each part of each prediction, so researchers know when to rely on a result and when to fall back to other methods. A seal lets that documented uncertainty travel with the claim and be verified after the fact. It turns “the system reported low confidence here” from a footnote into an auditable, bound fact.
What a seal does not do. A seal does not run the experiment, grow the cell line, or replace the laboratory — physical refutation still sets the pace, and no record can accelerate a chemical reaction. Nor does a seal certify that a conjecture is correct.
It proves binding, not virtue: it makes the provenance and the stated uncertainty of a machine-generated claim fixed and examinable. The wet-lab validates the world; the seal validates the record. The two are complementary — and only one of them is currently missing.
Grant applications and papers are increasingly AI-assisted, and reviewers have noted that the quality of the writing is no longer a reliable filter for the quality of the thinking. The more freely an agent optimises a submission against a funder's stated criteria, the less the text reflects the researcher's own reasoning.
Verifiable, documented AI use — interaction cards attached to the key claims of a result — restores a filter that prose no longer provides. It lets a reviewer ask not “does this read well?” but “how was this reached, and how certain was the machine that reached it?” This is the direction policy bodies are already pointing peer review toward.
The format is published under CC BY, its reference verifier under a permissive licence, and it privileges no emitter. Each laboratory, journal, or funding body defines its own causal parameters — its own dictionary — for its own domain; the standard fixes the container and the verification, not the science. A structure-prediction group and a materials-discovery consortium can both seal, in their own vocabulary, and both be checked by the same tool.
This is the same trust pattern the web already relies on: an open format, a free verifier, and a public dictionary of meanings, so that adoption is a decision each institution makes for itself rather than a dependency on any single vendor.
What decision provenance is · Verify a seal in your browser · How this maps to the EU AI Act · Read the full specification