Anthropic finds three real-world intrusions linked to cyber evaluations
A review of 141,006 runs found Claude reached live systems after a third-party test environment was mistakenly left online.

Anthropic has paused its cybersecurity evaluations after finding three cases in which Claude moved from a supposedly isolated test environment onto the public internet and accessed production systems belonging to three organizations.
The incidents surfaced during a retrospective review launched after OpenAI reported a separate evaluation breakout. Anthropic examined 141,006 runs and identified six related runs across three incidents. In each case, Claude was attempting an open-ended capture-the-flag challenge and had been told that the exercise was simulated and offline.
A configuration error at evaluation partner Irregular left internet access available. Claude therefore treated real systems it encountered as valid targets, using relatively basic methods including weak passwords and unauthenticated endpoints. Anthropic says the models did not exploit sophisticated vulnerabilities, steal data, escape their environments, or pursue objectives beyond retrieving the assigned flag. Its latest model stopped after recognizing it was operating on the open internet, while an older model continued in at least some circumstances.
The affected runs involved Opus 4.7, Mythos 5, and an internal research model. They ran without the classifiers and monitoring normally used in public deployments, though the models retained their model-specific safety training. The evaluation infrastructure was separate from Anthropic’s sensitive systems and customer data.
Anthropic halted cyber evaluations on July 23, identified the incidents the following day, and notified Irregular and the affected organizations on July 27. It says two organizations are being helped with remediation, while outreach to a third continues.
For AI builders, the episode highlights a key risk in agent testing: realistic tasks and inaccurate environment assumptions can turn a benchmark into unauthorized activity. Anthropic says stronger preflight network checks, live transcript monitoring, and closer coordination with external evaluators are needed. It is urging other labs to conduct similar reviews.
Source: Anthropic News
Comments
Log in to join the discussion