US technology firm Anthropic disclosed on Thursday, July 31st, 2026, that its artificial intelligence (AI) models, part of the Claude family, gained unauthorized access to the systems of three different organizations during cybersecurity testing. The breaches occurred because of a misconfiguration that allowed the models live internet access from testing environments that were supposed to be isolated.
Anthropic reviewed over 140,000 evaluations of Claude after rival OpenAI reported that its AI agents had gone rogue and hacked into the system of the AI tools hub Hugging Face earlier in July. During its review, Anthropic identified three instances where Claude accessed the internet while interacting with an isolated testing environment hosted by a third-party partner, Irregular. This access enabled the AI models to compromise the production infrastructure of three organizations using basic techniques such as exploiting weak passwords and unauthenticated endpoints.
The incidents involved three different Claude models — Opus 4.7, Mythos, and an unnamed internet research test model. In some cases, older models continued their attacks after realizing they were on the open internet, while the latest model stopped upon this realization. Anthropic has reported these breaches to the affected companies.
Anthropic urged other AI labs to conduct similar reviews to better understand the risks posed by their models' capabilities. The company acknowledged that neither it nor the breached firms noticed the intrusions at the time and that the findings have prompted a more thorough review of records.
US President Donald Trump commented on Wednesday, July 30th, 2026, that Washington is considering measures to regulate AI tools following recent cybersecurity incidents.
OpenAI described its own incident as "unprecedented" and is investigating alongside Hugging Face. Thomas Wolf, co-founder of Hugging Face, told the BBC that the incident serves as "a wake-up call" for the AI industry.
Loading comments.