Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good
Summary
White House science advisor Michael Kratsios alleged that Moonshot's Kimi K3 LLM was developed by illicitly distilling Anthropic's Fable. He also claimed Moonshot used export-banned Grace Blackwell 300s chips. Kratsios cited "covert industrial distillation" and "watermarks" of U.S. models. However, AI experts are skeptical that distillation alone explains Kimi K3's advanced capabilities. Fable became public July 1st, making a two-week turnaround for a strong model via distillation improbable. Researchers like Braden Hancock and Nathan Lambert suggest distillation's impact lessens as models near the frontier. Reinforcement learning, requiring significant infrastructure, is often needed. While Anthropic previously accused Moonshot of distillation, experts also note the technical expertise of Chinese teams. The allegations further involve Moonshot's access to banned Nvidia GB300 chips and servers in Thailand. This prompts calls for "know your customer" laws for data centers.
Key takeaway
For AI Policy Makers evaluating IP protection and export controls, this highlights challenges in distinguishing innovation from illicit transfer. You should prioritize robust "know your customer" regulations for global data centers and chip exporters. This prevents access to banned hardware. Also, consider distillation's diminishing returns as a primary threat. Focus instead on the broader technical capabilities of international teams.
Key insights
Experts doubt distillation alone explains advanced LLM capabilities, especially for frontier models developed rapidly.
Principles
- Distillation's impact diminishes as models near the frontier.
- Advanced LLM capabilities often require reinforcement learning.
- Chinese AI teams possess significant technical expertise.
Method
Distillation involves systematically querying a target LLM to generate data for post-training, sometimes asking for chain-of-thought or using prompts/responses for supervised fine-tuning (SFT).
In practice
- Monitor API usage for "deliberate capability extraction" patterns.
- Implement "know your customer" laws for data centers.
Topics
- LLM Distillation
- AI Export Controls
- Reinforcement Learning
- Kimi K3
- Anthropic Fable
- NVIDIA Grace Blackwell
- Geopolitics of AI
Best for: CTO, VP of Engineering/Data, Executive, AI Scientist, Policy Maker, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AI News & Artificial Intelligence | TechCrunch.