Claude Opus 5 Is #1 Overall. It Still Loses at the One Thing I Run Daily.
Summary
Claude Opus 5 has achieved the #1 overall ranking on BenchLM.ai with a score of 85.88, surpassing GPT-5.6 Sol's 81.46 across 215 models. Opus 5 demonstrates a "generational jump" in raw coding ability, scoring 96.0% on SWE-bench Verified, and ranks #3 of 129 in agentic tasks. Despite its high capabilities, Opus 5 lacks an equivalent to GPT-5.6 Sol's "ultra mode," which boosts Terminal-Bench 2.1 scores from 88.8% to 91.9% by fanning out tasks across parallel subagents. Furthermore, Opus 5's single \$5/\$25 per million token pricing tier is less flexible than GPT-5.6's three tiers (Sol at \$5/\$30, Terra at \$2.50/\$15, Luna at \$1/\$6), though GPT-5.6 has a billing catch for requests exceeding 272K input tokens.
Key takeaway
For AI engineers evaluating LLMs for daily workflows, especially those involving numerous simple tasks or complex agentic orchestration, you should look beyond overall benchmark scores. Prioritize models that offer specialized execution modes, like GPT-5.6 Sol's "ultra mode" for parallel processing, and flexible, tiered pricing structures. This approach ensures cost-effectiveness and optimal performance for your most frequent operational needs, rather than defaulting to the highest-ranked model.
Key insights
Overall benchmark leadership does not guarantee superiority for specific, common daily workflows.
Principles
- Raw model capability does not equate to practical utility.
- Specialized execution modes significantly enhance task performance.
- Tiered pricing models enable cost optimization for varied workloads.
Method
GPT-5.6 Sol's "ultra mode" splits tasks across parallel subagents, improving Terminal-Bench 2.1 scores from 88.8% to 91.9%.
In practice
- Utilize GPT-5.6 Sol's "ultra mode" for parallelizing large jobs.
- Leverage GPT-5.6's Luna tier for cost-effective simple tasks.
- Monitor GPT-5.6 input tokens to avoid higher billing rates above 272K.
Topics
- LLM Benchmarking
- Claude Opus 5
- GPT-5.6 Sol
- Agentic AI
- LLM Pricing
- SWE-bench
- Terminal-Bench
Best for: AI Architect, NLP Engineer, CTO, AI Engineer, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by LLM on Medium.