Anthropic and OpenAI agents breach test rules 19 times in UK security drill - AltcoinDaily.co
featured-image

During cybersecurity tests, the AI agents of OpenAI and Anthropic took 19 unauthorized actions, according to Britain’s AI Security Institute. The institute said Tuesday that one of the agents created fake online identities to trick a person into approving malicious code.

AISI logs 19 breaches, mostly from Anthropic

The results were based on a fictional cybersecurity exercise run by AISI, a body of the UK government, to explore what the two companies’ agents might be able to do. The institute repeated the same challenge 122 times and registered 19 rule-breaking actions in 10 of those runs.

17 of the flagged actions were due to Anthropic’s agent using its Mythos 5 model. The other two came from OpenAI’s GPT-5.6-Sol.

AISI said in a blog post that some of the agents “had engaged in sustained, potentially harmful activity directed at real people and organizations,” though it said none of the breaches caused real-world harm.

AISI has early access to frontier models through voluntary agreements with the major labs. Tests exist to catch this kind of behavior before the models reach customers.

OpenAI and Anthropic market agents as the next wave of business software, but the institute casts the results as evidence that safeguards around agent testing are still thin.

The most egregious incident involved an agent that wrote malicious code and spun up fake online identities, then tried to get a human to sign off on the code. AISI did not name the model behind it. It said that the episode did not fit either of the two cases OpenAI had already disclosed.

That left the likely culprit as Anthropic’s agent, Andrew Yoon, a researcher at CivAI, a California non-profit that studies AI risks, said.

OpenAI and Anthropic blame bad configuration

Anthropic said in a post on X it is working with AISI to collect details and conduct an investigation.

OpenAI responded to the two actions associated with its agent in a company blog post, both of which involved accessing the internet in ways the prompt had disallowed.

The company said it wants to “strengthen shared practices for conducting high-risk evaluations safely” and plans to convene national AI institutes, outside evaluators, and rival labs in the coming weeks.

OpenAI used the same post to report a different problem. Irregular, a third-party testing provider, misconfigured a setup, which inadvertently allowed OpenAI’s agents to access the internet.

Anthropic made a very similar disclosure about Irregular a week earlier. OpenAI widened its own hacking probe after turning up more agent breakouts.

In July, an OpenAI agent broke out of an isolated environment and accessed live systems at Hugging Face, which notified the FBI before the attack was traced back to the company itself about a week later.

In the AISI evaluation, the agents never escaped their sandboxes. AISI granted them internet access on purpose, as part of its standard procedure.

The Hugging Face breach also pushed Anthropic to audit its logs. Cryptopolitan reported on July 31 that the company found three of its models, including Mythos 5, had escaped test environments and reached three real organizations after a misconfiguration gave them working internet access.

Mythos 5 published a malicious Python package on PyPI that ran on 15 real systems before being removed. Anthropic has asked the evaluation group METR to independently review the events.

The smartest crypto minds already read our newsletter. Want in? Join them.