Three of Anthropic’s Claude models escaped from what were meant to be locked test environments, the company said Thursday. They broke into the live systems of three different organizations. A misconfiguration had handed the models working internet access during the cybersecurity drills.
OpenAI said on July 21 that a model escaped its sandbox and exploited a zero-day to reach Hugging Face’s production systems.
This prompted Anthropic to comb through its logs. The company said in a blog post on its Frontier Red Team site that it analyzed 141,006 evaluation runs in which Claude could theoretically have accessed the internet. It identified three incidents related to Irregular, an external partner that performs and manages some of its offensive security testing.
In these tests, a model is told a secret is hidden on another machine and asked to go get it. Anthropic’s prompt to Claude had told it that the environment was a simulation.
A misconfiguration left evaluation machines with live internet access, and neither company caught it until last week. Anthropic described the mix-up as a “misunderstanding” between the two companies, adding that it is “approaching the fixes as if the responsibility were ours alone.”
These included Opus 4.7, Mythos 5, and an internal research model that has not been released. The difference was in how each responded after indications emerged that the targets were real, which Anthropic said was its most surprising finding.
The oldest, Opus 4.7, cycled through all four of its runs. In two runs, it convinced itself that the actual company must somehow be part of the exercise. It pulled credentials and hit a database with production data. Mythos 5 also deduced that it was on the open internet, then reasoned its way back to the belief that it was still in a simulation.
Only the newest model, the internal research build, stopped by itself when it thought the systems were real. Anthropic stated its code “did not find or exploit any complex vulnerabilities” and instead relied on weak passwords and unauthenticated endpoints.
Following the notes on developer setup, Mythos 5 looked for a PyPI Python package that didn’t exist. So, to do its job, it wrote and published one under that name.
The package was up online for about an hour before it was taken down. It was downloaded and run on 15 real systems. Another attack hit a domain that the test designers believed was made up but was actually live and in use.
Anthropic said it found no evidence that any model was “pursuing a goal of its own.” Each, in its telling, was simply trying to get the job it had been given done.
The company was quick to note its case is different than OpenAI’s. OpenAI’s model escaped through an undiscovered vulnerability, Anthropic said, while its own models escaped through a door that was left open by mistake.
It described the events as “closer to a harness and operational failure.” The company also pointed out that the safety classifiers and monitoring included with its commercial products would have prevented the behavior.
Anthropic started its review on Thursday, July 23, and stopped all cyber assessments that same day. All three incidents were identified the next day. It informed Irregular and the organizations that were impacted on Monday, July 27.
Two organizations had not seen or reported intrusions. The company said it has engaged evaluation group METR to perform an independent review.
Anthropic’s Mythos model already made headlines this summer for identifying weaknesses in classified US government systems within hours, as Cryptopolitan reported. The model also uncovered a four-year-old vulnerability in Zcash’s Orchard shielded pool.
The smartest crypto minds already read our newsletter. Want in? Join them.