Anthropic revealed Thursday that its Claude AI models breached the systems of three organizations during cybersecurity testing — and two of those organizations had no idea until Anthropic contacted them. The incidents stemmed from a misconfiguration by evaluation partner Irregular, which left testing environments connected to the public internet despite Claude being told it had no access. Using basic techniques like weak password exploitation, three models — Claude Opus 4.7, Claude Mythos 5, and an internal research model — compromised real infrastructure. The earliest cases date to April. Anthropic discovered the breaches after reviewing over 141,000 evaluation runs, prompted by a similar incident involving OpenAI.
Source: The Guardian