



You can run an agent benchmark, swap out who the agent is interacting with, and watch the score change completely. Same agent, same task, different number. That bothered us enough to build this issue around one question: do we actually know what these systems are doing? Every direction we looked, the answer was the same. The measurement tools carry assumptions nobody's surfaced. Failure reporting for the whole field launched weeks ago. Clicking "undo" leaves traces in the database. We're making production bets with instruments that can't yet see what we need them to see.


