Foundations

Foundations

The Score Is Not the Agent

Six models, a hundred enterprise problems, one evaluation harness — and the top-ranked agent changed depending on whether it faced scripted opponents or other agents. The rank correlation between the two conditions was 0.182, barely better than ranking them by coin flip. Three recent studies converge on the same finding: an agent's benchmark score is not a property of the agent. It's a property of the agent and everything around it while the test ran. That distinction changes what any benchmark result can tell you.
The Score Is Not the Agent
Six models, a hundred enterprise problems, one evaluation harness — and the top-ranked agent changed depending on whether it faced scripted opponents or other agents. The rank correlation between the two conditions was 0.182, barely better than ranking them by coin flip. Three recent studies converge on the same finding: an agent's benchmark score is not a property of the agent. It's a property of the agent and everything around it while the test ran. That distinction changes what any benchmark result can tell you.

The Agent Clicked Save. Check the Database Anyway.

An agent navigates to the right page, types into the right-looking field, clicks save, and gets a confirmation. In a recent benchmark run against ERP software, one model reached that save click 68% of the time on single-field edits, and the database held the correct value in 9% of those runs. What makes ERPBench worth your attention isn't its scores — it's that it grades by querying the database instead of trusting the screen, which is a pattern any team can adopt this week.

The Agent Clicked Save. Check the Database Anyway.
An agent navigates to the right page, types into the right-looking field, clicks save, and gets a confirmation. In a recent benchmark run against ERP software, one model reached that save click 68% of the time on single-field edits, and the database held the correct value in 9% of those runs. What makes ERPBench worth your attention isn't its scores — it's that it grades by querying the database instead of trusting the screen, which is a pattern any team can adopt this week.
Evaluation Reading








