The word "production" does not appear in Melissa Pan's NeurIPS 2025 paper. Pan works at UC Berkeley's Sky Computing Lab, a systems group whose project list runs to agent runtimes, distributed infrastructure, evaluation frameworks. The default instrument there is code and the default evidence is a trace. That paper, co-led by Cemri, Pan, and Yang, worked through 1,642 execution traces (step-by-step records of what each agent did) from seven open-source frameworks for coordinating multiple AI agents, running benchmark tasks in programming, math, and general agent work. Out of it came a taxonomy of failure, sorted by whether the break sits in coordination between agents, in verification of the work, or in the design of the system itself. The limitation the authors flag is that the taxonomy may not be complete. They don't flag where the data came from.
By the following summer, the ICML 2026 paper Pan co-led carried 25 authors across five institutions, up from 13. It interviews 20 deployment teams and surveys 306 practitioners about proprietary agent systems running in production. No one annotated a trace; the team ran interviews, sorted the transcripts into categories by hand, and sent out a 47-question survey. The problem statement moved too: production agents are "typically proprietary," the paper says, and little public information exists about how they get built.
The ICML paper does not cite the NeurIPS paper. I went through the full text and the bibliography looking for the title, the framework name (MAST), the co-authors, the arXiv number. Not there. Six authors in common and the two papers don't point at each other. (I wrote about Pan's trajectory in issue #41; this is the one detail from her published record I keep turning over.)
I'm not going to guess why. Absences are weak evidence on their own, as anyone who has debugged from a log that stopped writing can tell you: "nothing happened" and "nothing got recorded" are not the same finding. What the gap does establish is narrow. The authors didn't present the second study as an answer to limits they hit in the first. The NeurIPS paper's stated boundary is taxonomy coverage. The ICML paper's is access to proprietary systems.
Lay them side by side anyway. The first studied how multi-agent systems fail by running open-source frameworks against benchmarks and reading what came back. The second studied how production systems work by asking the people who built them. What counts as data changed somewhere in between. Benchmark traces are reproducible, inspectable, bounded; you can rerun the thing. Practitioner accounts of closed deployments are none of that, and the environments being described resist reproduction for reasons that go past closed source code: live users behaving like live users, integrations with systems the team doesn't control, model versions changing underneath them. The paper is upfront about the cost, presenting its results as "qualitative evidence of practitioner approaches rather than as fixed prevalence estimates" and naming direct auditing of live systems as the work that should come next.
There's a trade in that. A survey fixes its questions before it goes out, so you get answers to the 47 things someone thought of in the spring and nothing on what you needed to know in the fall. Same bargain as a dashboard built before the outage.
If you're evaluating agent systems inside your own organization, the trajectory has a practical edge. Whatever moved the method, the constraint the ICML paper describes — the systems worth learning from are the ones you can't get inside — is probably standing in front of your evaluation process too. You won't see it until you ask a question your current setup can't answer.
- What practitioners actually constrain: Pan's ICML paper found that 68% of deployed agent systems executed fewer than 10 steps before human intervention, with teams systematically trading capability for controllability through sandboxes, restricted APIs, and short execution horizons.
- Monitoring at fleet scale: Anthropic disclosed that its monitors processed over a billion agent decisions in August 2026, blocking roughly one in 47,000 and reducing ~100,000 weekly flags to about 50 human-reviewed cases — a concrete picture of how observation becomes triage.
- Ordinary interfaces steering agents: The peer-reviewed DECEPTICON study found that agents were steered toward unwanted outcomes by dark patterns in over 70% of tasks, compared with 31% for human participants, even when task completion metrics looked fine.
- NIST on monitoring gaps: A March 2026 NIST report identifies fragmented logging, costly human review, and uncertain measurement validity as continuing challenges for deployed AI monitoring — the institutional side of the instrument problem Pan's work surfaces.

