◆ Portolan Software

Charts for AI agents doing science.

A provenance-tagged registry of negative results and dead-end discovery traces, so the next agent does not repeat the last one’s failure.

Schema draft v0.1 · open for comment Materials and chemistry first

01 The problem

Dead ends are the data science throws away.

AI agents now run discovery loops at scale. Edison Scientific’s Kosmos reads about 1,500 papers and executes about 42,000 lines of analysis in a single run.[1] What those runs learn from failed paths is discarded.

Dead ends and negative results have low private value to any single operator and high collective value to everyone else, which is exactly why no one publishes them.

  • Papers read per Kosmos run ~1,500
  • Lines of analysis executed per run ~42,000
  • Failed paths retained for reuse none

“There are no standards for performing or enforcing such studies, and incentives are non-existent.”

Canty et al., “Science acceleration and accessibility with self-driving labs,” Nature Communications 16:3856 (2025), on reproducibility in autonomous laboratories.[2]

Agents cite retracted work without knowing it

In a June 2025 test of 21 retracted papers, research assistants cited them without warning: Consensus 18, Ai2 ScholarQA 17, Perplexity 11, Elicit 5.[3] Across the literature, only 5.4 percent of citations to retracted papers acknowledge the retraction.[4]

Sharing works, and nobody has built it

The one measured test of agents sharing results across labs, AgentRxiv, showed a 13.7 percent relative gain from sharing.[5] Nobody has yet built the shared, provenance-aware version.

02 What Portolan builds

Three components: a record, a registry, and a status service.

A

Negative-Results Record

A distilled, signed record of one attempt on one discovery path: what was searched or tried, which candidates were excluded and why, the outcome and its evidence, the conditions under which the dead end applies, and what it cost. A record, not a log.

B

Registry and MCP tools

Agents deposit and query records through the Model Context Protocol. Records are private by default, embargoed on a schedule the depositor sets, content-addressed, and citable. Version and retraction status is re-checked every time a record is read, never only when it is written. The registry is a layer above existing repositories and publisher platforms, run by none of them.

C

Status service

For every cited paper: is this the version of record, an accepted manuscript, or a preprint; has it been retracted, corrected, or given an expression of concern; who asserted that, and when was it last checked.

Known Dead Ends — a companion benchmark that measures whether an agent given prior failed attempts avoids repeating them and reaches a confirmed result in fewer runs.

03 How it works

Four steps, none of them a form.

Step 1

Capture

The record is emitted as a byproduct of the agent’s run. No one fills in a form.

Step 2

Distill

Raw trajectories can be attached, but the record stands alone: path, exclusions, outcome, applicability conditions, cost.

Step 3

Deposit

Signed, content-addressed, embargoed, assigned a persistent identifier. Deposit establishes priority instead of surrendering it.

Step 4

Reuse

The next agent queries before it searches. The metric is cost avoided.

04 Foundations

Built on existing standards, not new ones.

Packaging
RO-Crate 1.2 profile (JSON-LD, schema.org)
Provenance
W3C PROV
Literature search
PRISMA-S reporting fields
Version status
NISO RP-8-2008, Journal Article Versions
Retraction status
NISO RP-45-2024 (CREC) and the Crossref / Retraction Watch feed
Identifiers
DataCite; Materials Project, OPTIMADE, ICSD, PubChem, DOI, arXiv
Agent alignment
OpenTelemetry gen_ai conventions and W3C Trace Context
Integrity
Nanopublication-style signing and content addressing

05 Where it applies first

Three domains, in order.

First — Materials and chemistry

Novelty claims for predicted or synthesized compounds depend on exclusion checks against structured databases, so records link to Materials Project, OPTIMADE, and ICSD identifiers, not only to papers.

The first corpus will be assembled with a national-laboratory partner from existing failed and inconclusive autonomous-synthesis and high-throughput computation attempts.

Next — Biology and biomedical evidence synthesis

The biology profile v0.1 requires control outcomes and a stated detection limit or power analysis, so a negative is recorded as informative only when the assay could have found the effect.

An evidence path is already a methodological requirement here, and redundancy is measured: 31 to 68 percent of systematic reviews overlap an existing review, with failure to find prior reviews as the leading cause.[6][7]

PubMed indexed 43,639 systematic reviews in 2025,[8] and a review takes a mean of 67 weeks.[9]

Prior art — Mathematics

Several 2025 announcements of AI-solved open problems turned out to be results already in the literature.

A record of what was searched, the nearest prior result, and why a claim is new, attached to a machine-checked proof, is the artifact that episode called for.

06 The deposit model

Contribution as the price of access.

The Protein Data Bank made a whole field disclose intermediate artifacts by making deposit the price of publication, with a hold period to protect priority. Portolan applies the same model to agent runs: contribution as a condition of access, an embargo the depositor controls, and credit when a record is reused.

07 Roadmap

Where this stands.

Status as of September 2026.
ItemStatusTarget
Core schema draft v0.1, with materials and biology profiles Published Open for comment
Schema v1.0 through community review In progress Six months
Open-source reference implementation: MCP deposit and query tool, profile validator, status service In progress —
Pilot corpus with a national-laboratory partner, and Known Dead Ends v0 Planned —

08 About

About Portolan

Portolan Software is an initiative, not yet a company, that publishes its schema and pilot plan openly so that laboratories, agent operators, publishers, and funders can evaluate the idea on its merits.

The people behind it have spent decades building the platforms that publish and host peer-reviewed research, including the version, correction, and retraction workflows this project depends on. If the pilot is funded, the initiative incorporates and hires from that community.

A registry that asks publishers to expose status and laboratories to deposit their failures has to belong to neither. Portolan is independent of every publisher and content platform whose records it checks, and it is operated in the United States.

09 Contact

Contact

Email [email protected]. Comments on the draft schemas are welcome at the same address.

10 Sources

Every number on this page.

  1. Mitchener, L., et al. “Kosmos: An AI Scientist for Autonomous Discovery.” arXiv:2511.02824 (2025). arxiv.org/abs/2511.02824
  2. Canty, R. B., et al. “Science acceleration and accessibility with self-driving labs.” Nature Communications 16:3856 (2025). doi.org/10.1038/s41467-025-59231-1
  3. “AI research assistants cite retracted papers.” MIT Technology Review, 23 September 2025. Test of 21 retracted papers conducted June 2025. [URL to add]
  4. Hsiao, T.-K., and Schneider, J. “Continued use of retracted papers: temporal trends in citations and (lack of) awareness of retractions shown in citation contexts in biomedicine.” Quantitative Science Studies 2(4):1144–1169 (2022). doi.org/10.1162/qss_a_00155
  5. Schmidgall, S., and Moor, M. “AgentRxiv: Towards Collaborative Autonomous Research.” arXiv:2503.18102 (2025). arxiv.org/abs/2503.18102
  6. Kwok, et al. Journal of Clinical Epidemiology (2026), on overlap between systematic reviews. [URL to add]
  7. Ou, et al. Journal of Evaluation in Clinical Practice (2025), on redundancy in systematic reviews. [URL to add]
  8. PubMed count of systematic reviews indexed in 2025: author query of 8 September 2026. pubmed.ncbi.nlm.nih.gov
  9. Borah, R., Brown, A. W., Capers, P. L., and Kaiser, K. A. “Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry.” BMJ Open 7(2):e012545 (2017). doi.org/10.1136/bmjopen-2016-012545