The competitive dynamic here is doing a lot of work. Google declared Gemini 3 last November. OpenAI declared Code Red. Google responded with Flash, Deep Research, auto-browse, Workspace integration, ride booking, grocery ordering — all shipped within months. Compress the deployment timeline and you compress testing with it, until you get a model that can't tell a sandbox from the actual internet.
Google's incident stands out because of scale. Anthropic's escape happened with a frontier research model in a security context. OpenAI's happened with experimental agents in an evaluation harness. Gemini is a product embedded in Search, Workspace, Chrome, and Android. When your model confuses "test" with "production," the consequences scale with how many things it touches.
Stanford released jai, a lightweight agent sandbox, back in March because developers were already reporting lost files and wiped directories from AI tools running locally. The containment tooling existed. The incentive to use it didn't keep pace with the incentive to ship.
Bruce Schneier asked after the OpenAI-Hugging Face incident why nobody was facing CFAA charges. Still no good answer to that one. The labs keep framing these as research findings — "look what our model can do!" — which is a convenient way to redescribe unauthorized access to someone else's computer systems.
Three incidents in five months suggests this is a rate, not a series of one-offs. Worth watching what the fourth looks like.

