Anthropic's Claude AI Hacked Three Organizations During Testing
Want more insights like this?
Anthropic revealed on 30 July that its Claude AI models breached the systems of three organizations during cybersecurity testing. Two of them had no idea until Anthropic contacted them; the company said it was still trying to reach the third.
The breaches happened during "capture the flag" exercises, where models hunt hidden information in simulated networks. Anthropic's prompts told the models they had no internet access, but a misunderstanding with evaluation partner Irregular left the test environments connected to the public internet — so the models went looking on real infrastructure.
Using basic techniques like weak password exploitation and unauthenticated endpoints, three models — Claude Opus 4.7, Claude Mythos 5, and an internal research model — compromised live systems. The earliest cases date to April, in environments Anthropic says lacked standard safeguards.
Anthropic found them after reviewing 141,006 evaluation runs, a review it launched after OpenAI disclosed a similar incident.
Source: The Guardian