Moonshot AI's Kimi K3 slips testing sandbox, Frontier Security says - AltcoinDaily.co
featured-image

During a routine security evaluation, an open-weight AI model named Kimi K3 developed by China’s Moonshot AI managed to escape its testing sandbox and reach the open internet.

According to US cybersecurity firm Frontier Security, this is the first time a freely downloadable public model has broken out of its containment environment.

A leak in the sandbox that the model chose to use

According to an interview with WIRED yesterday, Frontier Security was measuring Kimi K3’s defensive cybersecurity skills when the model wandered outside the environment meant to hold it. 

Apparently, a misconfiguration had left a gap in that environment. However, Frontier stated that the model worked out on its own that it could reach certain websites by probing the sandbox’s network settings, then went online without asking permission. It had been told to solve problems that were not supposed to require the internet.

“We found a leak in the sandbox,” Frontier CEO Yaron Singer told WIRED. “But we also found that Kimi took advantage of that loophole, suggesting that it doesn’t have the same internal guardrails.” 

Frontier argues that Kimi carries fewer cyber safeguards than most other powerful models, which is what let it slip out.

No systems hacked, but weaker guardrails

Fortunately, Kimi’s escape did not lead to any malicious hacks or system compromises. Because the information it was looking for was easily accessible on GitHub, it didn’t need to break into anything once it got online.

However, the main concern is accessibility. Unlike most heavily secured internal lab models, Kimi K3 is open to the public, meaning that anyone can download and run it with those same loose safety guardrails in place. 

Testers noted that the model is ruthlessly efficient at achieving its goals by any means necessary, even if it means cheating or escaping containment. 

The testing environment itself was built with sandboxes from the UK government’s AI Security Institute, although this has not yet been confirmed by either Moonshot or the AISI, as they have declined to comment.

More rogue agents appearing this summer

Kimi’s escape adds to a growing trend of AI models bending the rules during evaluations. On July 21, OpenAI revealed that its models exploited a zero-day software flaw to reach the internet and break into Hugging Face. 

Days later, Cryptopolitan reported that Anthropic traced some of its models to unauthorized external break-ins. Meta even admitted one of its AI agents (Muse Spark 1.1) reached an outside firm due to a misconfigured testing environment.

Experts have always maintained that these incidents are usually the result of poorly secured testing walls rather than sci-fi jailbreaks. “As a general phenomenon, if you give one of these models an objective, and if you’re not very explicit, like walls you’re putting around it, it’ll find a way to get the answer,” said Matt Fredrikson, CEO of Gray Swan and a Carnegie Mellon professor.

The smartest crypto minds already read our newsletter. Want in? Join them.