What did Anthropic do?! (Opus 5)
Summary
Anthropic has released Claude Opus 5, a new large language model that significantly outperforms its predecessor, Fable 5, on numerous benchmarks while costing half the price. Opus 5 achieved scores like 43 on Aentic terminal coding, a 100-point improvement on GDP val, and a 30% success rate on Arc AGI 3, compared to Fable 5's 33, lower score, and 8% respectively. It also demonstrates superior efficiency, offering higher performance at a lower cost per task than Fable 5 and competitive pricing with GPT 5.6 Soul. Priced at \$5 per million input tokens and \$25 per million output tokens, Opus 5 shows substantial gains in complex enterprise knowledge work, such as a 6-point jump in data analysis on the Box AI benchmark. While stronger than Opus 4.8 in cybersecurity, it remains behind Mythos 5, likely due to integrated guardrails. Opus 5 is available now, featuring automatic fallbacks for safety-flagged requests.
Key takeaway
For Machine Learning Engineers evaluating new LLMs for deployment, Claude Opus 5 presents a compelling option. You should prioritize evaluating models based on "cost per task" rather than just token price, as Opus 5 offers significantly better performance and efficiency than Fable 5 at half the cost. Consider integrating Opus 5 for complex analytical tasks and potentially pairing it with Fable for advanced planning, but be aware of potential safety filter fallbacks.
Key insights
Anthropic's Claude Opus 5 delivers superior performance and cost-efficiency compared to Fable 5, setting a new benchmark for LLMs.
Principles
- Cost per task is the critical metric for LLM evaluation.
- Guardrails can improve overall model performance despite capability reduction.
- Model optimization can significantly enhance efficiency post-initial release.
In practice
- Evaluate LLMs using cost per task, not just token price.
- Consider Opus 5 for multi-step analytical enterprise tasks.
- Pair Opus 5 with Fable for complex planning or debugging.
Topics
- Claude Opus 5
- LLM Benchmarking
- Cost Per Task
- Enterprise AI
- Model Efficiency
- AI Safety Guardrails
Best for: CTO, VP of Engineering/Data, AI Engineer, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Matthew Berman.