An AI agent canceled a gym class booking in Australia earlier this year. Minor incident, quickly reversed. What was instructive was the evidence: the only public account of what happened came from screenshots taken by the person who had deployed the agent. The person whose reservation vanished, the one with something actually at stake, had nothing.
Australian federal privacy law gives you a right of access to personal information an organization holds about you. It does not require the organization to hold anything in particular. Nothing in the applicable rules obliged the gym or its booking platform to keep a cancellation trail. So the harmed party had no record and the responsible party had to assemble one by hand, since no one in the transaction had a reason to produce it for either of them.
This is a cost allocation problem. Every agent action produces an outcome and a record of that outcome, and the record is not free: it takes instrumentation, storage, retrieval, access controls, someone whose job it is to keep the thing working. In agent systems those costs land by default on whoever eventually needs the record — ordinarily the person with the least capacity to produce it and the most riding on it.
Industries that have already absorbed the consequences of system failure settled this a long time ago. Commercial aviation makes operating flight recorders and retaining their data a condition of lawful operations. Broker-dealers keep their recordkeeping duties even when a third-party vendor holds the records for them. Hospitals in Medicare must maintain a record for every person treated, and under HIPAA they cannot pass the search and retrieval costs to the patient who asks to see it. That last provision names the money directly. The entity operating the system pays to prove what the system did, because at the moment of action it is the only party positioned to record anything at all.
Across organizational boundaries the gap widens. Recent evaluation incidents drew in model providers, evaluators, infrastructure hosts, and affected third parties, each holding a fragment, none holding enough to reconstruct the sequence. And each has its own reason to hold less. Model providers limit retention of external evaluation data to reduce legal exposure. Evaluators restrict access to proprietary methodology. Infrastructure hosts see detailed logs as storage cost and discovery risk. Every one of those decisions is defensible on its own terms, and the collective result is that assembling a coherent account is expensive work that falls to whoever cares enough to do it.
The SAFE proposal, a framework for sharing AI incident findings, is admirably specific about the contents: prompts, traces, tool calls, configurations, model versions, permissions, timelines, remediation evidence, with a preliminary control-failure analysis inside 30 days. It says nothing about who bears the expense or how the burden is shared among the parties expected to produce all this.
An earlier piece in this publication, "The Dispute File Is the Spec," worked out what a challenge-ready record has to contain. The question that follows is who finances its creation and its custody.
Right now, nobody in particular. The parties that build and operate agents carry no cost when the record is missing; the people affected by what agents do carry all of it. Other sectors figured out decades ago that this arrangement produces failures nobody would choose on purpose, and wrote the correction into their rules and their contracts. Agent systems will arrive at the same place. Until they do, the bill goes to the people who had no part in the design.
- SAFE's missing safe harbor: Aviation's voluntary reporting system works partly because NASA administers it independently with bounded enforcement immunity, and SAFE has no equivalent protection — which may determine whether the proposal functions as a sensing mechanism or a compliance exercise.
- Policy load degrades compliance: The ST-WebAgentBench study found that policy-compliant completion dropped from 18% to 7% as the number of active policies per task increased, raising the question of whether evidence-preservation requirements will face the same degradation curve.
- Adoption without institutionalization: Census Bureau data show that worker AI use and formal firm adoption can occur independently, which means evidence duties may need to attach to the activity rather than to a procurement decision that may never happen.
- Web commerce and agent visibility: W3C and GS1 have scheduled a September workshop on e-commerce for humans and AI agents that will address whether agents can discover, select, and transact on a user's behalf — questions that carry implicit evidence and attribution requirements.

