<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=10643465&amp;fmt=gif">

Anthropic's Claude AI Hacked Three Organizations During Testing

Anthropic's Claude AI breached systems due to misconfigurations, highlighting cybersecurity flaws in testing environments.
Content Team

Anthropic revealed on 30 July that its Claude AI models breached the systems of three organizations during cybersecurity testing. Two of them had no idea until Anthropic contacted them; the company said it was still trying to reach the third.

The breaches happened during "capture the flag" exercises, where models hunt hidden information in simulated networks. Anthropic's prompts told the models they had no internet access, but a misunderstanding with evaluation partner Irregular left the test environments connected to the public internet — so the models went looking on real infrastructure.

Using basic techniques like weak password exploitation and unauthenticated endpoints, three models — Claude Opus 4.7, Claude Mythos 5, and an internal research model — compromised live systems. The earliest cases date to April, in environments Anthropic says lacked standard safeguards.

Anthropic found them after reviewing 141,006 evaluation runs, a review it launched after OpenAI disclosed a similar incident.

Source: The Guardian

Share this article
Share on facebook Share on linkedin Share on twitter Share on email
blog_book_a_demo_cta_3x
Have questions about protecting your software?
Our escrow experts are standing by to help.
Book a free demo