[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)

· Source: Latent.Space - Www.latent.space · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems, Emerging Technologies & Innovation · Depth: Expert, long

Summary

Thinking Machines Lab has introduced Inkling, a new open-weights multimodal foundation model family, licensed under Apache 2.0. The flagship Inkling model is a Mixture-of-Experts transformer featuring 975B total parameters with 41B active, supporting a 1M token context window. It was pretrained on 45 trillion tokens across text, images, audio, and video. A lighter variant, Inkling-Small, offers 276B total parameters with 12B active. Positioned as a customizable base model rather than a benchmark leader, Inkling scored 41 on the Intelligence Index, surpassing other U.S. open-weight models like Nemotron 3 Ultra (38). Its architecture includes hybrid/sliding-window attention, relative positional encoding, short convolution layers, and MoE with two shared experts. The release garnered extensive day-0 ecosystem support from major inference stacks like vLLM and Modal, with Tinker API pricing starting at \$1.87 per 1M input tokens for 64K context.

Key takeaway

For AI Engineers evaluating open-weight multimodal models, Inkling presents a compelling option as the leading U.S. Apache 2.0 release. Its 975B parameter MoE architecture, 1M token context, and native multimodal reasoning offer a robust foundation for custom applications. You should consider integrating Inkling for agentic workloads or fine-tuning, especially given its strong day-0 ecosystem support and focus on controllable reasoning rather than just benchmark scores.

Key insights

Inkling is an Apache 2.0 multimodal MoE model, prioritizing customization and efficiency over raw benchmark supremacy.

Principles

Method

Inkling employs a Mixture-of-Experts transformer with relative positional encoding and short convolution layers, trained from scratch on 45 trillion multimodal tokens.

In practice

Topics

Best for: NLP Engineer, Computer Vision Engineer, CTO, AI Scientist, Machine Learning Engineer, AI Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Latent.Space - Www.latent.space.