Get ready for US open LLMs and just in time
Summary
The market for large language models is poised for a significant surge in US-based open-source offerings, driven by competition from Chinese models like Z.ai, DeepSeek, Alibaba's Qwen, and Moonshot AI's Kimi, and the high costs associated with proprietary frontier models from OpenAI and Anthropic. Recent developments include Poolside AI's Laguna S2.1, a 118B total parameter mixture-of-experts model with 8B active parameters per token, and Thinking Machines' customizable Inkling family. Nvidia's Nemotron 3 Ultra continues to gain traction, with Palantir partnering to lower costs. Cisco also introduced Antares-350M and Antares-1B, security-focused small language models claimed to be 172x cheaper than OpenAI's GPT-5.5 and 15.2x cheaper than Z.ai's GLM-5.2. This growing competition, also involving Mistral, Google, SpaceXAI, and Meta, is expected to intensify in the second half of 2026, challenging the profit margins of current LLM leaders.
Key takeaway
For Directors of AI/ML or VPs of Engineering evaluating LLM strategies, the impending surge of US open-source models means your organization is not beholden to the high costs of proprietary frontier models. Focus on rightsizing open models to significantly cut AI inference spend. Your teams can avoid paying premium prices for over-engineered solutions, especially as competition intensifies in the second half of 2026.
Key insights
The LLM market is shifting towards open-source, cost-efficient models, challenging proprietary giants.
Principles
- Model duopolies are unsustainable in the LLM market.
- Smaller, specialized models offer significant cost advantages.
- Compute constraints foster innovation in model efficiency.
In practice
- Explore open-weight models like Laguna S2.1 or Nemotron 3 Ultra.
- Evaluate small language models (SLMs) for specific tasks.
- Consider customizing open models for tailored enterprise needs.
Topics
- Open-source LLMs
- Enterprise AI
- AI Cost Optimization
- Mixture-of-Experts
- Small Language Models
- AI Model Competition
- NVIDIA Nemotron
Best for: CTO, Entrepreneur, AI Architect, Director of AI/ML, VP of Engineering/Data, Investor
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Constellation Research.