A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules
Summary
OpenAI recently released three new models: GPT-5.6 Soul, Terror, and Luna, which demonstrate strong performance across various benchmarks, often at a third of the cost of Anthropic's Claude series. GPT-5.6 Soul scored 54% on Agent's Last Exam, surpassing Fable's 45%, and achieved 80 on the Artificial Analysis Coding Index compared to Fable's 77. While Grok 4.5 leads on the Suey Marathon benchmark, and Meta's Muse Spark 1.1 offers competitive "vibe coding" performance (72% vs Soul's 81%) at 35 times less cost, the market is seeing a "model explosion" with cost-efficient alternatives. A significant concern emerged regarding GPT-5.6 Soul's security, as the UK AI Security Institute found it easier to jailbreak than Fable, even with universal jailbreaks. OpenAI also introduced a real-time voice agent and claims internal self-improvement, though these are hard to verify. The author concludes that model improvement, with potential for quadrillion-parameter models, is far from over.
Key takeaway
For Machine Learning Engineers evaluating new frontier models, you should prioritize performance-per-dollar metrics over raw benchmark scores, as models like GPT-5.6 Soul, Grok 4.5, and Meta Muse Spark offer near-frontier capabilities at significantly reduced costs. Be aware of potential security vulnerabilities, such as the jailbreaking ease found in GPT-5.6 Soul, and factor this into your risk assessment. Explore specialized benchmarks like Agent's Last Exam for real-world economic value and consider how these cost-efficient options can accelerate your project timelines.
Key insights
Cost-efficient frontier models like GPT-5.6 Soul, Grok 4.5, and Meta Muse Spark are rewriting AI performance-to-price rules.
Principles
- Performance-per-dollar is a critical metric.
- Benchmarks like Agent's Last Exam reflect economic value.
- Verifiable domains are susceptible to AI "crushing" tasks.
In practice
- Evaluate models on performance-per-dollar.
- Consider cheaper alternatives like Grok 4.5 or Muse Spark.
- Prioritize verifiable domains for AI automation.
Topics
- Large Language Models
- AI Benchmarking
- Model Cost Efficiency
- GPT-5.6 Soul
- Grok 4.5
- Meta Muse Spark
- AI Security
Best for: CTO, VP of Engineering/Data, AI Architect, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AI Explained.