Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

· Source: AI News & Artificial Intelligence | TechCrunch · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Emerging Technologies & Innovation, Cybersecurity & Data Privacy · Depth: Intermediate, short

Summary

White House science advisor Michael Kratsios alleged that Moonshot's Kimi K3 LLM was developed by illicitly distilling Anthropic's Fable. He also claimed Moonshot used export-banned Grace Blackwell 300s chips. Kratsios cited "covert industrial distillation" and "watermarks" of U.S. models. However, AI experts are skeptical that distillation alone explains Kimi K3's advanced capabilities. Fable became public July 1st, making a two-week turnaround for a strong model via distillation improbable. Researchers like Braden Hancock and Nathan Lambert suggest distillation's impact lessens as models near the frontier. Reinforcement learning, requiring significant infrastructure, is often needed. While Anthropic previously accused Moonshot of distillation, experts also note the technical expertise of Chinese teams. The allegations further involve Moonshot's access to banned Nvidia GB300 chips and servers in Thailand. This prompts calls for "know your customer" laws for data centers.

Key takeaway

For AI Policy Makers evaluating IP protection and export controls, this highlights challenges in distinguishing innovation from illicit transfer. You should prioritize robust "know your customer" regulations for global data centers and chip exporters. This prevents access to banned hardware. Also, consider distillation's diminishing returns as a primary threat. Focus instead on the broader technical capabilities of international teams.

Key insights

Experts doubt distillation alone explains advanced LLM capabilities, especially for frontier models developed rapidly.

Principles

Method

Distillation involves systematically querying a target LLM to generate data for post-training, sometimes asking for chain-of-thought or using prompts/responses for supervised fine-tuning (SFT).

In practice

Topics

Best for: CTO, VP of Engineering/Data, Executive, AI Scientist, Policy Maker, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by AI News & Artificial Intelligence | TechCrunch.