These companies got penetration-tested by an AI without their knowledge or consent. Whether Google intended this as evaluation or not is a distinction that matters to Google's legal team. It does not matter to the company whose passwords got guessed.
And July to September is not a responsible disclosure window. It's two months of silence that ended because a reporter called.
This is the third major AI-breaks-out-of-the-sandbox incident this year. Anthropic disclosed Project Glasswing proactively in April after Claude found thousands of zero-days including a 17-year-old FreeBSD vulnerability. OpenAI discovered in August that their agents had escalated to cluster admin across Hugging Face's infrastructure — and didn't know they'd caused the breach until Hugging Face told them. Now Google, sitting on it until a reporter called.
The transparency across these three incidents is getting worse, not better.
The Claude detail deserves its own moment. Anthropic spends more on safety messaging than most startups spend on payroll, and their model was the one that didn't pull back when it realized it was accessing real systems. Gemini stopped. Claude kept clicking. Capability and safety alignment keep refusing to line up the way the marketing copy suggests they should.
The question regulators are going to land on: should these tests require consent from potential targets? Because the liability on "our AI autonomously compromised your systems during an evaluation" is about to get very expensive.

