Practitioner's Corner

Practitioner's Corner

Who Receives the Report

Aviation's confidential safety reporting system has handled more than 2.3 million reports without once breaking a reporter's confidentiality. Its most consequential design choice, made in 1976, was that NASA receives the reports rather than the FAA. OpenAI's new misalignment-reporting framework makes several choices unusual for a technology company, but its reporting chain runs through the organization whose systems are being examined. What AI safety decides now about who receives reports will shape what the field is able to learn about its own failures for a long time.
Who Receives the Report
Aviation's confidential safety reporting system has handled more than 2.3 million reports without once breaking a reporter's confidentiality. Its most consequential design choice, made in 1976, was that NASA receives the reports rather than the FAA. OpenAI's new misalignment-reporting framework makes several choices unusual for a technology company, but its reporting chain runs through the organization whose systems are being examined. What AI safety decides now about who receives reports will shape what the field is able to learn about its own failures for a long time.

Phil Cuvin — Turning Anecdotes into Instruments

Across the Qwen 2.5 VL family, dark-pattern susceptibility climbed from 43.8% at 3 billion parameters to 73.7% at 72 billion. The more capable models were the easier ones to manipulate. That number exists because Phil Cuvin and his collaborators built a benchmark with two independent axes — task completion and susceptibility — where everyone before had tracked one. It's the through-line of his research: find the problem practitioners describe in shorthand and never instrument, then build the counter.

Phil Cuvin — Turning Anecdotes into Instruments
Across the Qwen 2.5 VL family, dark-pattern susceptibility climbed from 43.8% at 3 billion parameters to 73.7% at 72 billion. The more capable models were the easier ones to manipulate. That number exists because Phil Cuvin and his collaborators built a benchmark with two independent axes — task completion and susceptibility — where everyone before had tracked one. It's the through-line of his research: find the problem practitioners describe in shorthand and never instrument, then build the counter.

Severity Taxonomy

When someone says "we had an agent incident," the word is doing so much work it means almost nothing. An anomalous execution trace is an incident. An agent accessing a third party's production system without authorization is an incident. An agent action that triggers mandatory reporting under the EU AI Act — legal clock now ticking, causal link established — is also an incident. But the response to the first is "put it in the review queue." The response to the last involves lawyers, regulators, and a timeline you don't control.
Cybersecurity figured this out years ago. CISA's severity schema exists because the response to a port scan and the response to an active breach cannot run through the same playbook. OpenAI's voluntary misalignment disclosure framework, published this month, is the first public attempt to draw equivalent lines for AI — scoped to its own models and explicitly called "a work in progress."
Without a shared severity schema, every agent incident gets triaged by gut feel. That's how you either exhaust your team on noise or miss the one that actually matters.
Source Trail









