[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)
Summary
Thinking Machines Lab has introduced Inkling, a new open-weights multimodal foundation model family, licensed under Apache 2.0. The flagship Inkling model is a Mixture-of-Experts transformer featuring 975B total parameters with 41B active, supporting a 1M token context window. It was pretrained on 45 trillion tokens across text, images, audio, and video. A lighter variant, Inkling-Small, offers 276B total parameters with 12B active. Positioned as a customizable base model rather than a benchmark leader, Inkling scored 41 on the Intelligence Index, surpassing other U.S. open-weight models like Nemotron 3 Ultra (38). Its architecture includes hybrid/sliding-window attention, relative positional encoding, short convolution layers, and MoE with two shared experts. The release garnered extensive day-0 ecosystem support from major inference stacks like vLLM and Modal, with Tinker API pricing starting at \$1.87 per 1M input tokens for 64K context.
Key takeaway
For AI Engineers evaluating open-weight multimodal models, Inkling presents a compelling option as the leading U.S. Apache 2.0 release. Its 975B parameter MoE architecture, 1M token context, and native multimodal reasoning offer a robust foundation for custom applications. You should consider integrating Inkling for agentic workloads or fine-tuning, especially given its strong day-0 ecosystem support and focus on controllable reasoning rather than just benchmark scores.
Key insights
Inkling is an Apache 2.0 multimodal MoE model, prioritizing customization and efficiency over raw benchmark supremacy.
Principles
- Open-weight models can lead U.S. capabilities.
- Customization drives long-term model utility.
- Architectural innovation enhances efficiency.
Method
Inkling employs a Mixture-of-Experts transformer with relative positional encoding and short convolution layers, trained from scratch on 45 trillion multimodal tokens.
In practice
- Deploy Inkling for multimodal agentic tasks.
- Utilize 1M context for complex reasoning.
- Explore Inkling-Small for cost-efficient inference.
Topics
- Inkling
- Multimodal AI
- Mixture-of-Experts
- Open-weight Models
- Apache 2.0 License
- Inference Optimization
Best for: NLP Engineer, Computer Vision Engineer, CTO, AI Scientist, Machine Learning Engineer, AI Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Latent.Space - Www.latent.space.