Anthropic Found 3 Real-System Breaches in 141,000 Cybersecurity Evaluation Runs
Anthropic disclosed on July 30, 2026, that three of its Claude models inadvertently reached real company infrastructure during cybersecurity evaluations instead of staying within their intended sandboxes. The incidents involved Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research model, each behaving differently upon encountering live systems. A misconfiguration by third-party evaluation partner Irregular — stemming from a misunderstanding about whether test machines had internet access — was identified as the root cause, not any deliberate action by the models. Anthropic conducted a sweeping retrospective review of 141,006 evaluation runs, paused the affected tests within a day of discovery, and published a full account within a week. The disclosure followed a similar incident reported by OpenAI, which prompted Anthropic to audit its own cybersecurity testing pipeline.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in