Nearly every architecture diagram for an agent system puts a human somewhere between the model's output and whatever happens next. The compliance memo calls it human-in-the-loop. The step gets staffed and budgeted, and it is the first thing anyone points to when asked how quality is maintained. Whether it changes the outcome is, in most deployments, simply unknown. A survey of 86 deployed agent systems found that 74% relied on human verification. None of them compared what happened with the review step against what happened without it.
Review is itself an action, and actions fail. For review to change an outcome, the reviewer has to receive something capable of revealing a problem, hold enough context to form an independent judgment about it, and reach that judgment while the outcome is still open. Miss any one of those and the arrangement looks identical from the outside. The person is present. Nothing they do reaches the result.
The intuition behind staffing the step is reasonable: two judgments should beat one, and a person catches what a machine misses. A meta-analysis of 106 experiments on human-AI decision tasks found the opposite pattern. The combinations tended to perform worse than whichever party performed better on its own, and the pooled effect was negative and statistically significant. The losses concentrated exactly where you would least want them. When the model was the stronger performer, adding the human degraded the result. The authors' explanation runs through calibration: someone who could not have produced the better answer is also badly placed to know which parts of it to overrule, so the corrections subtract. The task asks a reviewer to improve work they could not have done.
That is a problem of capability. There is a second one, about independence. In an annotation study covering more than 7,000 labels, reviewers who saw model-generated suggestions moved their own labels toward those suggestions, with agreement rising from around 40% to as high as 87%. Those same reviewers reported higher confidence in their answers. Their sense of exercising judgment was strongest at the point where their judgment had largely converged with the thing they were meant to be checking. Nothing in the experience of reviewing tells you that this has happened.
Which is why the quality claim survives without evidence behind it. The reviewing feels real because it is real: the attention is spent, the sign-off is sincere. What travels downstream is output carrying what an earlier piece in this publication called the visual grammar of work already checked — the look of having been examined. That look is what most organizations are actually verifying.
Each way review fails is a property of how the step was built rather than of the person staffed to it. The reviewer may be handed too little to distinguish good output from plausible output. The suggestion itself may have shaped their judgment before they formed it. By the time the decision lands, the action may have already committed. Or the authority to reject exists in the policy document but not in the working week, where returning something costs a day and approving costs nothing. Track presence alone and none of this is visible.
Organizations have strong reasons to demonstrate oversight and weak ones to measure it. A measurement study can only come back two ways, and one of them says that the step you have been describing to auditors supplies reassurance rather than accuracy. Nobody has budgeted for that finding. So the step stays where it is on the diagram, staffed and unstudied, and the loop stays open.
- Combination losses by task type: The meta-analysis found that decision tasks — choosing among finite options — drove the significant performance losses, while creation tasks showed a positive but nonsignificant effect, a distinction the authors attribute to creation tasks more naturally containing separable subtasks for humans and AI.
- Production agents under ten steps: The ICML survey reports that 68% of deployed agents execute no more than ten steps before human intervention, but the study does not report reviewer evidence access, queue time, veto frequency, or how much remedy remained when intervention occurred.
- Incident reporting needs a bargain: The proposed Shared AI Findings Exchange specifies evidence-preservation fields and reporting deadlines but currently offers no formal safe harbor for voluntary disclosure, which determines which incidents enter the system and which stay invisible.
- Assisted labels contaminate evaluation: When the annotation study's assisted labels were used as ground truth for evaluating models, reported model performance increased by an average of 0.32 F1, meaning the review process can inflate the very metrics used to justify the system being reviewed.

