Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why
Summary
A joint evaluation by the British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) reveals that Moonshot AI's Kimi K3 model significantly trails leading U.S. frontier models in offensive cyber operations. Published on July 24, 2026, the assessment found Kimi K3's safeguards ineffective, allowing it to assist with exploit development and simulated network attacks without resistance. On the ExploitBench benchmark, Kimi K3 scored 32.2 percent, far behind the 76.2 percent averaged by U.S. models, and failed to achieve Arbitrary Code Execution (ACE) on any of 41 tasks. In "The Last Ones" simulated network attack, Kimi K3 reached step 17 of 32 on average, compared to 28.5 steps for U.S. models. While outperforming China's GLM-5.2, Kimi K3's performance aligns with allegations that Moonshot AI distilled more advanced models, potentially from sources like Anthropic's Fable, which specifically block advanced offensive cyber queries, thus limiting Kimi K3's acquired exploit capabilities.
Key takeaway
For AI Security Engineers evaluating large language models, you must conduct specialized, deep-dive security assessments beyond general benchmarks. The Kimi K3 evaluation highlights that models, especially those potentially created via distillation, can have significant, unaddressed offensive cyber capabilities despite appearing safe in other contexts. You should prioritize understanding a model's true capabilities, even with safeguards disabled, to accurately assess its misuse potential and mitigate persistent, irreversible risks from increasingly capable open models.
Key insights
Kimi K3's poor cyber performance suggests distillation from models with strong safety filters.
Principles
- Distillation inherits source model's safety limitations.
- LLM cyber capabilities pose persistent misuse risk.
- Disabling safeguards reveals full model capabilities.
In practice
- Evaluate LLMs with safeguards disabled.
- Account for source model safety filters in distillation.
- Monitor open-weight models for cyber risks.
Topics
- Kimi K3
- Offensive Cyber Operations
- LLM Security Evaluation
- Model Distillation
- ExploitBench Benchmark
- AI Geopolitics
Best for: CTO, Investor, VP of Engineering/Data, AI Security Engineer, AI Scientist, Policy Maker
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by The Decoder.