The worst incidents I've ever worked had nothing broken in them. Nothing was down. The request completed inside its latency budget, the error rate never moved, every dashboard insisted the system was healthy, and a customer was on the phone getting the wrong answer and was entirely right to be furious about it. There is no page in the runbook for that.
Agent commerce is being built to fail in exactly that shape, and not because anyone is acting badly. It's a measurement problem.
The IMF's April note on agentic payments draws a line I keep coming back to. On one side, intent formation: the probabilistic part, where a model interprets what you meant, searches, compares, decides. On the other, authorization and settlement: deterministic rule-checking against limits you set, like amount, merchant category, timing, identity. An instruction that clears those checks becomes eligible for execution, and execution does not pause to reinterpret it. The note is careful to present this as an analytical frame rather than a description of today's rails. It also says, without ornament, that agents may misread intent or favor a provider's incentives over the user's welfare.
Everything checkable sits on the deterministic side: retained payment instructions, order confirmations, identity attestation, dispute reason codes. The side that determines whether you were served well emits almost nothing you can measure, contest, or bill against. That is not a logging gap. The comparison that would settle the question never entered the record, because only one option was ever acted on.
Fidelity belongs beside task completion and safety as a product property. Of the three, it's the only one with no instrument that survives contact with a real transaction.
Hold that while I go through what exists.
Real early work is underway. ShopperBench scores how faithfully an agent holds to a shopper persona; APOLLO tests whether tool-using agents recover and apply stated and implied preferences, where GPT-4o came in a shade over 51 percent. Both are careful and peer-reviewed, and both build the preference in advance so there's something to score against. Neither watches a live buyer and asks what they would have chosen from alternatives they never saw. That's roughly the boundary NIST's draft guidance marks when it notes that benchmarks suit tasks with discrete, verifiable answers while subjective open-ended ones generally need field testing.
The dispute rails have the same blind spot. Mastercard's current chargeback guide reaches further than most people assume, covering non-delivery, failed refunds, and goods that are defective, misdescribed, or unsuitable for their purpose. What it does not reach is the case agent selection will generate at volume: a properly authorized purchase, delivered as described, working fine, and worse than the option the agent never surfaced. No reason code fits. In "The Delegation Muscle" we argued that payments already own the machinery for authorization and dispute. They do. That machinery establishes who held authority. It says nothing about whether the choice was any good.
Regulated markets crack this occasionally. The SEC's finding that a broker's order routing had cost customers $34.1 million produced a number only because a public quoted price stood there to subtract against. Product selection has no quote. Price, quality, delivery timing and taste all feed the decision, so there's no counterfactual sitting around waiting to be measured.
Ranking at least left the options you didn't take on the page, where you might notice them. Agent selection doesn't render them at all. The cost lands on the buyer as a substitution they can't see, inside a transaction where every other participant can demonstrate they performed correctly.
My guess is that the first credible attempt at pricing this won't be a benchmark. It'll be sampling real completed purchases and asking the humans what they'd have picked instead.

