Practitioner's Corner

Practitioner's Corner

Melissa Pan's Missing Citation

Melissa Pan co-led a NeurIPS 2025 paper analyzing how multi-agent systems fail under benchmark conditions, then co-led an ICML 2026 paper surveying practitioners about proprietary agent deployments in production. Six authors appear on both. The second paper does not cite the first. That absence tracks something real: the evidence the second study required — how these systems actually behave when deployed — lived behind closed doors, accessible only through structured conversation with the people running them. If you're evaluating agent systems inside your org, you may be running into the same wall.
Melissa Pan's Missing Citation
Melissa Pan co-led a NeurIPS 2025 paper analyzing how multi-agent systems fail under benchmark conditions, then co-led an ICML 2026 paper surveying practitioners about proprietary agent deployments in production. Six authors appear on both. The second paper does not cite the first. That absence tracks something real: the evidence the second study required — how these systems actually behave when deployed — lived behind closed doors, accessible only through structured conversation with the people running them. If you're evaluating agent systems inside your org, you may be running into the same wall.

You Can't Stage Someone Else's Website

A housing association in Wales built an automation for food-bank referrals. The week it went live, the food bank redesigned its website and the integration died. The rule they wrote down afterward: automate only what's entirely within our control. That reads as timidity until you trace what sits underneath it. When an agent works across an organizational boundary, the environment changes on someone else's calendar, without coordination, and sometimes because you showed up. No amount of rehearsal closes that gap, because the stage isn't yours.

You Can't Stage Someone Else's Website
A housing association in Wales built an automation for food-bank referrals. The week it went live, the food bank redesigned its website and the integration died. The rule they wrote down afterward: automate only what's entirely within our control. That reads as timidity until you trace what sits underneath it. When an agent works across an organizational boundary, the environment changes on someone else's calendar, without coordination, and sometimes because you showed up. No amount of rehearsal closes that gap, because the stage isn't yours.

Mechanism Sidebar

A September 22 preprint tested a deterministic enforcement gate — a component sitting between an AI agent and its execution environment, refusing any action that violates compiled task state. On an airline benchmark, it did exactly what you'd want: caught ineligible cancellations, blocked policy violations, raised the agent's pass rate from 0.39 to 0.54. "Does this action violate policy?" was answerable from database state alone, so the gate's judgment was reliably correct.
Then they ran the same architecture against a project-management benchmark, where the equivalent question — "is this task due right now?" — required interpreting free-text fields. The gate's matcher misread those cues often enough to refuse roughly 22 selections per run. Seventeen of those were actually correct. Performance dropped below the ungated baseline.
The gate didn't malfunction. It enforced with full consistency against a model of the task that didn't match the task. And consistency applied to a wrong model is how a safety component becomes a new failure source. The judgment baked into the gate at build time was specific to one environment's state representation. Move the environment, keep the gate, and you've automated confident wrongness.
Reading List








