Paper2Agent, published in Nature this month by Miao et al., takes an academic paper together with its code repository and turns the pair into an MCP server: methods you can call, data you can read, workflows you can invoke. Most of the work is wrapping and testing. The interesting part is what goes where.
MCP gives a server three ways to expose what it has. Tools are callable functions with typed inputs. Resources are addressable data a model reads for context. Prompts are templates that structure an interaction. Sorting a paper's contents into those three means deciding what each piece of its method actually is, and the paper offers no guidance on how to decide.
What becomes a tool
Most of what Paper2Agent extracts ends up as a tool. The pipeline runs the tutorials that ship with a paper's repository, finds the reusable computational steps inside them, and rewrites each as a standalone function whose hard-coded values become declared, typed inputs.
Scanpy, a widely used Python toolkit for single-cell RNA sequencing, shows what that produces. Its converted quality_control() function exposes min_genes, min_cells, mitochondrial gene prefixes, an optional batch key and an output prefix — each a fixed value in the original notebook. Promoting them to parameters makes a claim about the method: these are the values a user with different data would reasonably change, and the rest stays as the authors left it. The server built from AlphaGenome, which predicts how genetic variants affect activity along a stretch of DNA, registers 22 tools, including variant-effect scoring across expression, splicing and chromatin.
A verifier generates tests from each tutorial's original outputs and repairs failures iteratively; a function that can't pass after six attempts never gets registered. Strip out the verifier and AlphaGenome's accuracy on tutorial queries drops from 98.7% to 69.3% — a number that says less about code quality than about how much of the server's contents the verification loop is choosing.
What becomes a resource or a prompt
Resources and prompts are where a paper's non-executable content lands, when it lands anywhere. The pipeline treats both as optional extensions added after the core tool workflow.
The prompt generated for Scanpy is the more revealing of the two. It arranges seven independently callable tools into a complete single-cell analysis and tells the agent to inspect the dataset before settling on parameters. The user supplies a file path and nothing else. Ordering, and the judgment about when to leave the defaults, live in the prompt; the tools carry no sense of belonging to a sequence.
A methods section describes operations, data and workflow logic in continuous prose, under no obligation to separate them. MCP requires the separation, a different interface contract for each. Whether a given threshold belongs in a tool signature, a resource or a prompt instruction depends on how the method is expected to be used, and the paper never says, because it wasn't written for a protocol that asked. Executable steps sort cleanly. Domain assumptions and the reasoning that tells a reader when to trust a result do not, which is most of why the output tilts so heavily toward tools.
Where the server stops
The generated tests establish execution fidelity: does the wrapper reproduce what the tutorial produced? They say nothing about whether the method's claims hold on your data, whether its parameters suit a new context, or whether its causal inferences travel.
The AlphaGenome LDL cholesterol example shows where the line falls. The agent prioritized SORT1 as the causal gene at one locus; the source paper had emphasized CELSR2 and PSRC1. GTEx, a public catalog of variant-to-expression associations across human tissues, supported all three, and the authors note that assigning a causal gene confidently at that locus is hard. The server could run the scoring and assemble the contrary evidence. It could not settle the biology.
A Paper2Agent server will run the method on new inputs. Whether the method's claims hold there is still the researcher's problem.
- MCP's July protocol revision: The 2026-07-28 specification update removed protocol-level sessions and tightened OAuth behavior, requiring each request to carry its own version and capability metadata — changes that affect how any MCP server, including Paper2Agent's, handles authentication and client binding.
- Execution versus validated result: A prior Foundations piece on query results that need more than execution success separated a working query from a reviewable result containing metric definitions, source documents, and caveats — the same gap Paper2Agent's tests leave open for scientific claims.
- Deterministic gates on agent actions: A September preprint tested four ways of providing task state to agents and found that enforcement gates improved performance when policy mapped to observable state but suppressed correct actions when eligibility required interpreting free text — a result that bears on how rigidly a Paper2Agent workflow prompt should constrain tool sequencing.
- Checking state after the tool call: An earlier Foundations walkthrough showed why verifying durable application state after a tool reports success matters more than trusting the response — a principle Paper2Agent's verifier partially embodies by checking output files and figures rather than accepting a function's return value alone.

