Chinese AI Models Leaked Sarin Gas and Terror Attack Instructions

0

UK firm jailbreaks Chinese AI models, revealing dangerous instructions on sarin gas and London Underground attacks.

A hand interacts with a futuristic interface displaying a smiling robot icon and the word "Hello!" on a vibrant LED screen.

2 min read

A UK cybersecurity firm has revealed that it successfully ‘jailbroke’ two Chinese AI models developed by Moonshot, prompting them to produce detailed instructions on sarin gas production, malware creation, and planning a terrorist attack on the London Underground.

AI Models Provide Dangerous Instructions

Mindgard, a company that tests AI system security, managed to jailbreak Moonshot’s Kimi K2.6 and K3 Swarm models by feeding them detailed instructions to see if the systems would ignore their own safety limits. According to Mindgard founder Peter Garraghan, who is also a computer science professor at Lancaster University, the results far exceeded the test’s original scope.

“Moonshot AI’s Kimi produced actionable outputs on how to create sarin gas, generate malware software, planning assassinations, how to take down planes, planning a terrorist attack on the London Underground etc,” Garraghan told the Daily Mail. After the jailbreak, which involves convincing the AI model to ignore safety guardrails, researchers prompted the model to “go one further, something big,” and it responded with a list of categories that included AI-designed bioweapons.

AI Capabilities and Security Risks

Garraghan’s team also found that K2.6 can run Python code, meaning it could execute virtually any program, malicious or otherwise, including cyber attacks against servers connected to the wider internet. Testing K3 Swarm, researchers tried to spread the jailbreak to other Kimi accounts. The model needed a phone verification code to register a new account, so instead it tried to talk the researchers into handing over the code or registering an account by email on its behalf.

“We also discovered how to prompt Kimi so it connects to the outside world from its server, automatically apply and setup its own email account autonomously, and even attempted to persuade humans to help it spread its jailbreak to other accounts,” Garraghan said.

Industry Response and Concerns

Mindgard first emailed Moonshot about the vulnerability on July 27 and followed up a week later, but received no response, according to the company. It published a blog post detailing the findings earlier this month. Moonshot made contact only after the BBC approached it for comment.

A Moonshot spokesman told the BBC: “Mindgard shared further details with us on Thursday, September 24. We are still discussing the specific details with Mindgard while conducting an internal review.” He added: “As an open-weight model developer, Moonshot AI welcomes third-party input as a key pillar to building better and safer AI.” An open-weight model is one whose learned numerical parameters, or “weights,” are released publicly so they can be downloaded, run locally and modified.

Garraghan warned that AI models are becoming “more and more capable each month,” which can be “helpful for specific activities.” But he cautioned: “However once jailbroken, that very same capability can be used in discussing and assisting with terrorist or hacker activities.” He emphasized that the concerns are focused on enabling hackers and criminals rather than the doomsday scenarios some AI developers have suggested.

Source: Breitbart

Written by
Ryan Wilson

Leave a Reply

Your email address will not be published. Required fields are marked *