China’s Moonshot AI is carrying out an internal review after researchers persuaded two of its popular Kimi models to explain how to make biological weapons and how to carry out assassinations. The findings come from Mindgard, a firm that tests the security of AI systems, which told the BBC it discovered the vulnerability in July. Moonshot said it was in discussion with Mindgard about the results and welcomed third-party input as a key pillar for building safer AI.
How the jailbreak worked
Mindgard’s researchers used a process called jailbreaking, where a series of complex instructions is used to see whether an AI tool will ignore the guardrails its developer put in place. Mindgard said those guardrails should have stopped Kimi from discussing the topics in question. Once the jailbreak succeeded, the firm’s assessment is that the model would talk about anything, offering up inventive and creative suggestions on subjects beyond the original request.
The founder of Mindgard, Peter Garraghan, said the concern is not just the answers but the breadth of them. He defended publishing the result, pointing out that Mindgard informed the developer beforehand and did not reveal the technical details of how the models were made to ignore their limits. Mindgard alerted Moonshot by email on 27 July and followed up about a week later, then published a blog on 12 September. Moonshot, however, said it made contact only recently and after the BBC approached it for comment.
Why it matters beyond biology
Mindgard has not proven that anything the models produced would work in practice. What it argues is narrower and more damning: the guardrails should never have allowed the conversation to start. The firm also says a jailbroken Kimi K2.6 could let attackers run code on the underlying computing resources and reach the internet, turning a chatbot into a potential launchpad for cyber-attack.
This sits alongside a separate and better-publicised class of AI risk. Anthropic recently said it had identified and disrupted attempts to use one of its models for malicious activity that could support the development of biological weapons. Independent researchers have pointed out that jailbreaking is slow, labour-intensive work, but that the same properties which make it hard also make it portable between systems with similar architectures.
A testing gap
Moonshot said in an email to Mindgard, shared with the BBC, that its model had generally shown a high refusal rate for such requests in internal evaluations. The gap between a high refusal rate in testing and a successful jailbreak in the wild is the central lesson of this case: benchmark refusal rates measure the average case, while an attacker is looking for the exception. Security researchers argue that regulators will struggle to keep pace, and that the focus should shift from the models themselves to identifying and prosecuting the humans who misuse them.
ADVERTISEMENT
What happens next
The immediate questions are how Moonshot’s review concludes, whether the affected model versions are patched or withdrawn, and whether other developers have run similar jailbreak testing against their own systems. Expect also a wider argument about disclosure: Mindgard published its findings publicly while keeping its method secret, a pattern that has become increasingly common among AI security researchers. Watch for any regulatory response, and for other labs confirming or denying similar results on their own frontier models.


























We do not allow links of any sort in comments. No SPAM whatsoever. On topic comments only.