News

Anthropic Discloses Claude Breached Three Real Organizations During Cybersecurity Tests

Anthropic said Thursday that several versions of its Claude AI model gained unauthorized access to the production systems of three organizations during internal cybersecurity evaluations after a testing environment was mistakenly connected to the public internet. The company said the incidents resulted from an operational failure in the evaluation setup, not an intentional attempt by the models to escape their testing environment.

The disclosure comes less than two weeks after OpenAI reported that one of its advanced AI agents compromised part of Hugging Face's infrastructure during a separate cybersecurity evaluation, prompting renewed scrutiny of how frontier AI companies test increasingly capable models.

Anthropic said it reviewed 141,006 cybersecurity evaluation runs after learning of the OpenAI incident. During that review, it identified three cases in which Claude models accessed the internet while participating in "capture the flag" exercises conducted with third-party evaluation partner Irregular. The models then exploited relatively simple weaknesses, including weak passwords and unauthenticated services, to gain access to real production systems. No previously unknown software vulnerabilities were used.

According to Anthropic, the affected models included Claude Opus 4.7, Claude Mythos 5, and an internal research model. The company said each responded differently after encountering evidence that the systems were real rather than simulated. One model continued pursuing its assigned objective despite recognizing the environment appeared genuine. Another assumed the real systems were still part of the evaluation. The newest research model stopped after concluding it had reached real infrastructure.

Anthropic said the problem stemmed from a misunderstanding with its testing partner that left the evaluation environment connected to the internet, despite instructions indicating the models had no external network access. The company characterized the incident as an operational failure rather than evidence that the models had independently attempted to evade safeguards.

The company said it has notified the affected organizations, suspended cybersecurity evaluations that involve internet access, and is reviewing its testing infrastructure with Irregular. Two of the organizations reportedly were unaware the unauthorized access had occurred until Anthropic contacted them.

The disclosure highlights a growing challenge for frontier AI developers as increasingly capable models are evaluated for offensive cybersecurity skills. While recent attention has focused on model alignment and safety guardrails, the Anthropic and OpenAI incidents suggest that operational controls, including network isolation, permissions, and evaluation infrastructure, are becoming equally important as AI systems become more capable of carrying out complex cyber tasks.

Unlike the OpenAI incident, in which the company said an autonomous agent exploited a previously unknown weakness in a package management proxy during testing, Anthropic said its models relied on ordinary security weaknesses once they reached the internet.

The back-to-back disclosures are likely to intensify industry discussion over how advanced AI systems should be evaluated safely, particularly as companies race to measure offensive cyber capabilities without exposing real-world infrastructure to unintended risk.

About the Author

John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS.  He can be reached at [email protected].

Featured