



We kept getting stuck on one number. An agent scored 100% on correctness and 58% on safety in the same run. It completed the job and leaked credentials it had no business touching. Both scores are accurate. That's the whole problem: the metric you naturally reach for can be orthogonal to the property that actually matters. Builders sense this already. The ones who could run agents longer keep choosing shorter loops, because the receiving organization can't verify what happened fast enough. This issue is us working through what that gap means for people who have to ship.


