An agent describes a photograph to someone who cannot see it. The description is useful because the person has no independent access to the image. And because they have no independent access to the image, they cannot check the description. Both facts follow from the same condition.
The assistive case makes the structure easy to see, but it is not where the structure lives. It appears wherever an agent stands between a user and information the user cannot reach on their own. Someone who reads no Turkish gets a translation. A patient reads a generated summary of a specialist's notes. A browser agent visits a site the user has never opened, reads what is there, and comes back with an account of it. The user can confirm that a summary arrived. They can judge whether it hangs together, whether it sounds plausible, whether the sentences parse. What they cannot do is set it beside the page it came from, because if they could do that, they would not have sent the agent in the first place. The inability to verify is not a gap in training or a lapse in diligence. It is the reason the agent was invoked.
Economics named this category more than fifty years ago. In 1973, Michael Darby and Edi Karni described goods whose quality the buyer cannot evaluate even after consumption, and called them credence goods. Auto repair is the standard illustration. The car runs again, which tells you nothing about whether the mechanic replaced the part that was failing, or replaced a different part, or replaced nothing and billed for the labor. Some things you can inspect before you buy. Some you find out about by using them. Credence goods resist both, and in the deployments that matter most, agent output belongs in that third category. The user consumes it and is no closer to knowing whether it was faithful to the source.
In a study of conversations conducted through machine translation, fluent output concealed errors that users had no way to detect. More telling: when the system produced awkward phrasing, users sometimes read the awkwardness as something their conversation partner meant. A machine failure got reassigned to a human being's character. A 2026 study of poetry translation found that monolingual readers preferred GPT-4's renderings while bilingual readers preferred the human translator's. Access to the original changed the verdict. The people who could check disagreed with the people who could not.
The natural response is transparency: show the user how the agent arrived at its answer. A ZEW experiment on algorithmic advice complicates that. Explaining the algorithm's process helped participants recognize that bias was present, and made their decisions worse, because knowing bias exists tells you nothing about which direction it runs or how far. What improved performance was revealing the correct answer. Independent access to ground truth. Which is precisely the thing these users do not have, by definition rather than by oversight.
Plenty of an agent's output remains checkable. A payment cleared. A file downloaded. A paragraph is grammatical. Those are the mechanical parts, and they are checkable because they were never behind the access barrier. Whether the agent represented what it actually found is still on the far side of it.
This condition will hold for more and more of what agents are asked to do, because that is where the value is. An agent that fetches something you could have fetched yourself saves you an afternoon. An agent that reaches what you cannot reach is worth considerably more, and is considerably harder to govern, because the person with the strongest interest in the answer being right is the person least equipped to find out. We will build around this. Building around a condition is not the same as removing it.
- Bias recognition vs. correction: A ZEW experiment found that explaining an algorithm's process helped users detect bias while worsening their actual decisions, because recognizing a problem and sizing it are different competencies.
- Appeals as lossy sensors: An HHS Inspector General report found that Medicare Advantage beneficiaries appealed only 1% of denials while insurers overturned 75% of those that reached first-level review, a reminder that low challenge rates cannot be read as evidence of accurate initial decisions.
- Loyalty to which user: A W3C draft proposes that web user agents owe duties of protection, honesty, and loyalty to the user, but in enterprise settings the identity of the legitimate principal is itself an unresolved institutional question.
- Intent versus authorization: An IMF framework for agentic payments argues for separating probabilistic intent from deterministic settlement, acknowledging that a valid mandate can prove an agent stayed within spending limits without proving it faithfully represented what the user actually wanted.

