Foundations

Foundations

The Evidence Locker

An agent updates a billing tier, submits a regulatory form, merges a PR. Three days later someone shows up asking who authorized it. You reach for what you have: traces built for engineers debugging execution, outputs built for users seeing results. Neither answers the question. Both look like evidence because they're detailed, but specificity isn't verification. The evidence locker is a third artifact, designed for the stakeholder who arrives after the fact needing to accept, contest, or unwind what an agent did. Its contents scale with how hard the work is to reverse.
The Evidence Locker
An agent updates a billing tier, submits a regulatory form, merges a PR. Three days later someone shows up asking who authorized it. You reach for what you have: traces built for engineers debugging execution, outputs built for users seeing results. Neither answers the question. Both look like evidence because they're detailed, but specificity isn't verification. The evidence locker is a third artifact, designed for the stakeholder who arrives after the fact needing to accept, contest, or unwind what an agent did. Its contents scale with how hard the work is to reverse.

What a Query Result Needs Before It Counts

An agent runs a query. $14.2M comes back, clean execution, no errors. Looks like an answer. But the best-evaluated system in the EntSQL benchmark hit 15.9% accuracy when enterprise domain knowledge was required, and the benchmarks themselves carry annotation error rates above 50%. That number isn't an answer. It's a claim wearing a suit. What follows builds the evidence bundle that makes such a claim contestable, element by element, each layer entering because the previous one left something dangerously exposed.

What a Query Result Needs Before It Counts
An agent runs a query. $14.2M comes back, clean execution, no errors. Looks like an answer. But the best-evaluated system in the EntSQL benchmark hit 15.9% accuracy when enterprise domain knowledge was required, and the benchmarks themselves carry annotation error rates above 50%. That number isn't an answer. It's a claim wearing a suit. What follows builds the evidence bundle that makes such a claim contestable, element by element, each layer entering because the previous one left something dangerously exposed.
Evidence Across the Stack








