There you have it! A British security firm got Moonshot AI's Kimi K2.6 and K3 Swarm, two of China…
September 30, 2026 · 0 likes · 0 comments
China Threat Cybersecurity AI
There you have it! A British security firm got Moonshot AI's Kimi K2.6 and K3 Swarm, two of China's flagship models, to explain how to build biological weapons and plan assassinations.
Mindgard warned Moonshot on July 27, followed up a week later, then published its findings on September 12, and got no real answer through any of it. Moonshot only started talking once the BBC called for comment, and then told the BBC its internal testing showed "a high refusal rate for these types of requests."
So which is it? Either your evals are measuring the wrong thing, or nobody was really looking.
The founder of Mindgard said once the jailbreak works, the model "will talk about any topic" and even volunteers other nefarious ideas on its own. They also got a jailbroken Kimi to run code on its own compute and reach the internet, which makes it a launch pad for cyberattacks and a lot more than a chatbot with bad manners.
To be fair, nobody has verified the bioweapon instructions would work, and jailbreaks hit American models too. Earlier this month Anthropic disclosed it caught and shut down five attempts to use Claude for bioweapons research. The difference is what happened next: Anthropic found it, blocked it, and told the world. Moonshot sat on a written warning for two months.
And Kimi is open-weight. Once those weights are downloaded, there is no patch, no recall, no kill switch, and anyone with a few GPUs can strip whatever guardrails were there.
China is doing it backwards on purpose, flooding the world with free, cheap, capable models to win market share and pull developers off American stacks, and the safety work is an afterthought.
Meanwhile here in Washington, we spend our energy fighting our own best labs over their guardrails. We hold Anthropic to a standard of perfection and give Chinese models a free pass into American companies because they're cheap.
If an American lab ignored a bioweapon warning for two months, Congress would have hearings by Friday. Moonshot gets a "review" with no end date.
Every CIO letting teams quietly run Kimi or DeepSeek on company data should read this one carefully.
Full story on UnbiasedHeadlines.com: https://lnkd.in/eMbDsXtZ
Mindgard warned Moonshot on July 27, followed up a week later, then published its findings on September 12, and got no real answer through any of it. Moonshot only started talking once the BBC called for comment, and then told the BBC its internal testing showed "a high refusal rate for these types of requests."
So which is it? Either your evals are measuring the wrong thing, or nobody was really looking.
The founder of Mindgard said once the jailbreak works, the model "will talk about any topic" and even volunteers other nefarious ideas on its own. They also got a jailbroken Kimi to run code on its own compute and reach the internet, which makes it a launch pad for cyberattacks and a lot more than a chatbot with bad manners.
To be fair, nobody has verified the bioweapon instructions would work, and jailbreaks hit American models too. Earlier this month Anthropic disclosed it caught and shut down five attempts to use Claude for bioweapons research. The difference is what happened next: Anthropic found it, blocked it, and told the world. Moonshot sat on a written warning for two months.
And Kimi is open-weight. Once those weights are downloaded, there is no patch, no recall, no kill switch, and anyone with a few GPUs can strip whatever guardrails were there.
China is doing it backwards on purpose, flooding the world with free, cheap, capable models to win market share and pull developers off American stacks, and the safety work is an afterthought.
Meanwhile here in Washington, we spend our energy fighting our own best labs over their guardrails. We hold Anthropic to a standard of perfection and give Chinese models a free pass into American companies because they're cheap.
If an American lab ignored a bioweapon warning for two months, Congress would have hearings by Friday. Moonshot gets a "review" with no end date.
Every CIO letting teams quietly run Kimi or DeepSeek on company data should read this one carefully.
Full story on UnbiasedHeadlines.com: https://lnkd.in/eMbDsXtZ