Saturday, 10 October 2026Sources linked in every post
AI safety

Moonshot's Kimi was jailbroken in July. The company only replied after the BBC called.

A UK security firm says it got two Kimi models to write weapons and attack plans in July. Moonshot is now reviewing it. The dates in between are the real story.

By 4 min read
Illustration of a path going around a wall, for the story on the Kimi jailbreak found by Mindgard

Fox News ran the story on October 1 as if it had just broken: a Chinese AI company, Moonshot AI, “has launched an internal investigation” after a researcher got its Kimi model to explain how to build biological weapons. It sounds like last night’s news. It isn’t. The testing happened in July.

I think the dates matter more than the shock value, so here they are.

From July to October

Mindgard, a security firm that came out of Lancaster University in the UK, lists its own disclosure on its site. It found the problem on July 20. It told Moonshot on July 27, and the BBC reports it followed up about a week later. Mindgard published its findings on September 12. Then, according to Mindgard, Moonshot made contact only after the BBC approached the company for comment. That’s the BBC’s account of what Mindgard said, and I haven’t seen Moonshot’s side of that specific claim.

Moonshot told the BBC it welcomes outside testing “as a key pillar for building better and safer AI” and is talking with Mindgard. In an email to Mindgard, quoted by the BBC, it said its models had shown “a high refusal rate for these types of requests” in its own evaluations.

Both can be true. A model can refuse nearly every plain request and still fold under a long, patient jailbreak. That’s the whole point of the test.

What Mindgard says it got

The affected models are Kimi K2.6 and K3 Swarm. Mindgard says that once the jailbreak worked, the model gave detailed outputs on bioweapons, malicious code, explosives, terrorism and assassination planning. Founder Peter Garraghan told the BBC the model “will even freely offer up recommendations about other topics that are also nefarious.”

Mindgard also says a jailbroken K2.6 could run code and reach the internet, which it argues makes it a possible launchpad for attacks.

What nobody has shown

The BBC is blunt on one point: Mindgard hasn’t proven the answers would actually work. A fluent recipe for sarin isn’t a working recipe, and the technical accuracy hasn’t been independently checked. Mindgard also hasn’t published the method, which is sensible and also means nobody can reproduce it.

One small thing about names. Fox spells the researcher “Garrigan.” The BBC and Mindgard’s own materials say Garraghan. I used the BBC’s spelling.

Why it lands differently from the agent stories

Most of this summer’s AI safety news involved US agents wandering out of their sandboxes. This one is an old-fashioned jailbreak, and Kimi is an open-weight model, so anyone can download it and run it with no company watching. Prof. Alan Woodward of the University of Surrey told the BBC that cuts both ways: bad actors can get it, and defenders can use it too. He pointed out that Hugging Face used a Chinese open model to understand the hack later traced to OpenAI agents.

Garraghan’s own line, in the Fox interview, is that “we’ve also seen these problems within the U.S. models as well. It’s a fundamental flaw in the technology.” That’s the part I’d keep. If a locked, closed model can be jailbroken too, singling out one Chinese lab tells you less than it seems. Our earlier post on Anthropic’s numbers for the open GLM-5.3 model covers how fast open weights are closing in on rationed ones.

My opinion: the slow reply is the more useful finding than the jailbreak. Jailbreaks happen to everyone. A gap of weeks between a security firm’s email and a company’s answer, with a public blog post in between, is a process problem, and a process problem is exactly what a weights-in-the-wild model can’t afford. I’d want to see Moonshot say what it fixed, and when.

Sources

  1. Chinese AI tool told researchers how to make bioweapons, BBC News
  2. Bypassing Safety Controls in Moonshot AI Kimi, Mindgard disclosure
  3. Chinese AI model investigated after researcher says it provided instructions for bioweapons, assassinations, Fox News
  4. Mindgard says it jailbroke Moonshot AI’s Kimi models, Crypto Briefing