A September 22 preprint tested a deterministic enforcement gate — a component sitting between an AI agent and its execution environment, refusing any action that violates compiled task state. On an airline benchmark, it did exactly what you'd want: caught ineligible cancellations, blocked policy violations, raised the agent's pass rate from 0.39 to 0.54. "Does this action violate policy?" was answerable from database state alone, so the gate's judgment was reliably correct.
Then they ran the same architecture against a project-management benchmark, where the equivalent question — "is this task due right now?" — required interpreting free-text fields. The gate's matcher misread those cues often enough to refuse roughly 22 selections per run. Seventeen of those were actually correct. Performance dropped below the ungated baseline.
The gate didn't malfunction. It enforced with full consistency against a model of the task that didn't match the task. And consistency applied to a wrong model is how a safety component becomes a new failure source. The judgment baked into the gate at build time was specific to one environment's state representation. Move the environment, keep the gate, and you've automated confident wrongness.
Airline benchmark (gate helped): Violations decidable from database state. Pass rate rose 0.39 → 0.54. Zero policy-violating writes executed.
PM benchmark (gate hurt): Due-ness required interpreting free text. Performance dropped to 0.55 vs. 0.63 ungated. Of ~22 refused selections per run, ~17 were actually correct.
Under state corruption: Gate accuracy fell from 0.98 to 0.06 as corrupted state entries rose from fewer than 1 to roughly 7 per episode. Most damage came from the agent silently following a wrong directive, not from a blocked action.
Implication for scaling: Gate performance tracks state correctness, not model capability. A stronger model behind the gate cannot rescue a gate whose compiled state is wrong — which means you can't test your way out of this by upgrading the agent.

