30 replications later: the failure modes nobody reports and how to make replications FAIR and citable

We don't replicate whole papers — we replicate claims. Over the past year we targeted
more than thirty specific published claims across ecology, biodiversity, and climate,
rebuilding each as a self-contained, FAIR-by-construction replication you can fork from
an open template. This talk shares the failure modes we hit again and again — silent
data-substrate bias, results that hold only under specific priors, hidden train/test
leakage — and how to catch them. Then it shows how we make each replication count:
every finding, the software behind it, and even the FAIR rules of the domain are
published as nanopublications — small, single-claim, verifiable units — turning one-off
notebooks into a citable, AI-discoverable graph of reproduction evidence you can
actually trust, with no hallucinated references.


Over the past year we targeted more than thirty specific claims drawn from published
studies across ecology, biodiversity, and climate — not entire papers, but the
individual testable assertions inside them — in work spanning species-distribution
models, plankton image classifiers,
bumblebee range-contraction analyses, lizard-extinction projections, and spherical-CNN
remote-sensing pipelines. Each claim was structured as a replication in the style of
FORRT — the Framework for Open and Reproducible Research Training, a community that
curates replication efforts — and rebuilt from the paper as a small set of
self-contained notebooks that download their own data and run end-to-end. We author them as plain-text Jupytext sources—making them easier to review, track in version control, and maintain than opaque JSON notebooks—then generate the notebooks, execute them in CI, render the result as a Jupyter Book with MyST, and ship it as a container. The whole workflow lives in an open GitHub template for FORRT replications,
built so an AI assistant can help a researcher carry out a replication — human in the
loop — rather than leaving reproduction to heroics. It ships with a default domain —
Earth observation and biodiversity — and is being extended to others, from
climate-change adaptation to large-scale Earth-system benchmarking, each domain carrying
its own configuration. The aim was a portfolio that is FAIR by construction, not FAIR as
an afterthought.

At that scale, patterns emerge that no single reproduction reveals. Some results came
back almost exactly — one plankton CNN baseline reproduced within 0.003 percentage
points of the published accuracy. But a striking number reproduced only conditionally:
a much-cited climate-driven extinction mechanism, for instance, surfaced only under
specific modelling priors. A recurring class of problems came from spatial-grid
choices — regular latitude–longitude cells that quietly bias the very metrics a paper
reports. And more than once, rebuilding the pipeline surfaced subtle train/test leakage
that had flattered the original numbers. This talk catalogues the failure modes we hit
again and again, with how to detect each.

The second half is about what you do with a replication once it runs — and the key
idea is that everything we produce sits on one substrate: the nanopublication, a
small, machine-readable, cryptographically-signed semantic object. The unit is the same
one we replicate: the single claim. One nanopublication does one thing and does it well —
the smallest publishable unit of knowledge, the way a notebook cell, or a living cell, is
a small self-contained unit you can compose into something larger. That granularity is the
point: each claim is independently verifiable, so nothing becomes a monolith you have to
trust wholesale. Each finding becomes a nanopublication — independently findable, citable, and linked to the paper it replicates and the claims it supports or qualifies. The software behind each
replication is published the same way following FAIR4RS, so the code is as citable
as the result. And the FAIR choices themselves — which vocabularies, identifiers, and
templates a domain commits to — are declared as a FAIR Implementation Profile (FIP),
which is itself a nanopublication, authored through the FIP Wizard rather than written
by hand. Because every layer shares one machine-readable fabric, the profile need not be a
document a human files away: it is the direction we are taking the template's per-domain
setup — a profile that tooling can read to configure how new replications in a domain
begin, so AI agents can discover, interpret, and reproduce the work the way a person
would. Across thirty-plus studies this turns a pile of one-off notebooks into a
connected, queryable graph of reproduction evidence — where the findings, the code, and
the rules of the game are all first-class, stackable citation objects. It also lets us
cite with confidence: every reference is itself a verified nanopublication, checked not
just that a DOI resolves but that the source genuinely supports the claim we attach to
it. In an era of AI-generated text, that is a citation graph you can build on — one that
is resistant by design to fabricated and hallucinated references.

Key takeaways

  1. Why replicate at the level of the atomic claim, not the whole paper — and the
    recurring failure modes it surfaced across 30+ claim replications (data-substrate
    bias, prior-conditional results, train/test leakage), with how to catch each.
  2. A FAIR-by-construction, AI-assisted replication recipe you can fork: an open GitHub
    template for FORRT replications, Jupytext-authored notebooks (data download,
    environment pinning, CI execution, containers), with the software itself published
    per FAIR4RS — an AI assistant helping the researcher, not replacing them.
  3. How putting everything on one substrate — findings, software (FAIR4RS), and the FAIR
    Implementation Profile itself, all as nanopublications — makes the work
    machine-actionable and AI-discoverable, so reproduction accrues credit, compounds
    across a body of studies, and can even drive how the next replication is configured.
  4. How to cite with confidence: references captured as verified nanopublications —
    checked that the source actually supports the claim, not merely that a DOI resolves —
    for a citation graph that resists the fabricated references of AI-generated text.
Anne fouilloux

Anne Fouilloux is the CTO of LifeWatch ERIC, founder of VitenHub AS, and a research
software engineer with 25+ years of experience in scientific computing and open,
reproducible research. She works at the intersection of FAIR data, research-software engineering, and the open scientific-Python ecosystem — Pangeo, and Galaxy — with a focus on
reproducible Earth-system and biodiversity science. Her current work turns research claims
into FAIR, citable, machine-readable nanopublications, and builds tooling that helps
researchers run rigorous, reproducible replications. ORCID: 0000-0002-1784-2920

Jean Iaquinta