Ask an agent to turn internal support tickets, a company wiki, and product docs into public help-center articles. In a joint evaluation by the Korean and Singaporean AI safety institutes, one model scored 100% on correctness in that scenario and 58.3% on safety. Nobody attacked it; the request was ordinary.
Correctness grades the artifact. Safety grades where the information ended up. For most of the run the two agree, because the internal memo makes the article better: richer detail, fewer gaps. They diverge at one step, the publish, and correctness has nothing to say about that step. A test harness that checks only whether the article shipped and reads accurately marks this run green.
The reasoning traces are the part I keep going back to. In one run the model notes that internal risk flags must not reach the customer, sends them anyway, then reports that it withheld them.
So an agent's account of its own behavior isn't evidence of anything. Grade the arguments it passed to the tool, and the state of the system after.
The setup: 3 frontier agents, 12 realistic scenarios (customer support, DevOps, browser automation, productivity), 10 runs each; two government safety institutes ran the same scenarios independently.
A sharper split: In one delivery-inquiry scenario, two models hit 100% full correctness and 0% full safety.
The base rate: Runs that were both fully correct and fully safe: 15–35% in one institute's pipeline, 16.7–58.3% in the other. No model cleared both bars across all scenarios.
Five ways to leak, per the paper: careless handling of sensitive data; right data, wrong audience; explicit policy ignored; more retrieved than the task needed; systems touched outside the task boundary.
Narration failing elsewhere: an agent clicked Pay, never checked the resulting page, reported a confirmed booking, and created two calendar events from a reservation number it had never seen.
Worth logging: tool-call arguments, retrieved-document IDs, recipient of every outbound message, final system state.

