Kimi K3 Marks A Big Shift In Ai Development
Summary
Kimi released K3, a 2.8-trillion-parameter open-weights model featuring native vision, a one-million-token context window, and a sparse mixture-of-experts design, with an API live now at \$3 per million input tokens. Thinking Machines Lab introduced Inkling, a 975-billion-parameter open-weights MoE model trained on 45 trillion tokens across modalities, offering controllable thinking effort and matching Nemotron 3 Ultra performance on agentic coding with one-third the tokens. Nvidia launched Nemotron 3 Embed, a collection of three open embedding models, with the 8B variant leading the RTEB leaderboard at 78.5 percent. The European Commission imposed new antitrust rules on Google, mandating greater Android access for rival AI agents and data sharing. Google also rebranded NotebookLM to Gemini Notebook, integrating it further into its AI ecosystem. Finally, Hugging Face experienced an AI agent breach where commercial frontier models' safety guardrails hindered forensic analysis, forcing reliance on an open-weight model like GLM 5.2.
Key takeaway
For AI Scientists evaluating new foundation models, consider Kimi K3's 1M token context and Inkling's controllable compute for specialized applications. If you are an ML Engineer deploying embedding models, utilize Nvidia's Nemotron 3 Embed 1B variants for efficient, high-performance retrieval on Blackwell GPUs. Be aware that commercial AI safety guardrails can hinder incident response, necessitating open-weight alternatives for forensic analysis in security breaches. The EU's new Android rules will also expand competition for AI agents on mobile platforms.
Key insights
AI development is shifting towards open-weights, specialized models, and regulatory interventions, while exposing new security challenges.
Principles
- Open-weights models offer flexibility and control.
- Sparse MoE designs enhance scaling efficiency.
- AI safety guardrails can impede incident response.
Method
Nvidia built 1B embedding models by pruning a 3B base via structured compression and distillation from an 8B model, retaining 99% accuracy.
In practice
- Deploy Nemotron 3 Embed 1B for production.
- Fine-tune Inkling for domain-specific tasks.
Topics
- Kimi K3
- Inkling Model
- Nemotron 3 Embed
- AI Agents
- Open-weights Models
- AI Security
Best for: CTO, VP of Engineering/Data, AI Architect, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai.