A research paper carries its findings. It also carries the conditions under which those findings can be challenged. The methods section says what was done and why. The statistical assumptions are written down so a reviewer can argue with them. The limitations section exists because the authors decided to tell you where they think they might be wrong. None of that is formatting convention; it is what lets a second party check the work without trusting the people who did it.
Paper2Agent, published in Nature this month, turns research papers into callable servers: you point one at your own data and it runs the paper's methods on it. The system keeps more than you might expect: manuscripts, code, datasets, validation tests. But the interface is built for invocation. You call a tool, you get a result, and the result does not arrive with the paper's stated limitations, its statistical assumptions, or any mark separating what the research established from what the agent worked out on the spot. Some of that material is in there, reachable if you know to ask. Reachable-if-you-ask and visible-by-default are not the same condition.
The review surface narrows. A paper can be challenged line by line: why this method, whether that threshold was justified, what happens if you relax an assumption. An agent running those same methods gives you inputs and outputs, with the reasoning between them folded out of sight. The published peer review file for Paper2Agent has reviewers circling this exact problem, asking how a user would tell the generated agent's limitations from the ones it inherited from the source research.
The version question gets harder. Academic publishing has specific machinery for handling change. NISO defines a Version of Record as a fixed, citable artifact. Crossref recommends that substantive changes get their own identifiers and visible correction notices rather than silent substitution. Both are imperfect and largely voluntary, but they encode a principle: when knowledge changes, the change should be visible and linked to what it replaced.
A callable agent has no equivalent convention. It can update its tools, retrain a model, or drift because something upstream changed and nobody downstream was told. The FORCE11 Software Citation Principles ask you to cite exact versions and platforms, which assumes the service gives you something stable to cite: a commit hash, a tagged release, a snapshot. Where what's exposed is a live endpoint and nothing else, you can write down what you should have recorded. You cannot go back and get it.
The line between findings and interpolation blurs. Paper2Agent's authors report that their agents answered questions the source papers never addressed; in one case an agent prioritized a different gene than the original researchers had emphasized. That may well be useful. But the response does not mark where the paper's work stops and the agent's begins, and the stored material carries the authority of peer review, which extends to claims peer review never looked at.
The three run together because they come from one trade. A paper is a container built so its claims can be examined, and examination wants fixed text, stated assumptions, and a visible seam between what was established and what someone is inferring on top of it. A callable agent is a container built so its capabilities can be used, and use rewards fluency, integration, a clean answer. Moving from the first to the second is a real gain in usability. The cost side of it isn't displayed anywhere in the interface, and whether anyone attends to it is mostly a question of incentives: shipping capability is legible and rewarded, preserving someone's ability to argue with a result is neither, at least not until someone needs to have the argument.
- Evaluation shifts with environment: A recent preprint found that a deterministic enforcement gate raised one agent's benchmark score from 0.39 to 0.54 but pushed another below its baseline when eligibility depended on interpreting free-text cues — a concrete case where the instrument's environment changes what it measures.
- Memory provenance is adjacent: The problem of unmarked interpolation extends beyond callable papers to agent memory itself, where fluent output can conceal the boundary between retrieved material and fresh inference.
- Hidden human rescue contaminates claims: Reuters reported that a consumer agent was tested with contractors quietly handling some calls, raising reported success rates while obscuring what the automation actually accomplished on its own.
- Audit infrastructure before audit evidence: California's SB 813 and AB 1405 create verifier designation criteria and an auditor registry without yet establishing whether audited systems actually fail less often — institutional machinery that precedes the outcome data it will eventually need.

