Claude Opus 5 Is #1 Overall. It Still Loses at the One Thing I Run Daily.

· Source: LLM on Medium · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Software Development & Engineering · Depth: Intermediate, quick

Summary

Claude Opus 5 has achieved the #1 overall ranking on BenchLM.ai with a score of 85.88, surpassing GPT-5.6 Sol's 81.46 across 215 models. Opus 5 demonstrates a "generational jump" in raw coding ability, scoring 96.0% on SWE-bench Verified, and ranks #3 of 129 in agentic tasks. Despite its high capabilities, Opus 5 lacks an equivalent to GPT-5.6 Sol's "ultra mode," which boosts Terminal-Bench 2.1 scores from 88.8% to 91.9% by fanning out tasks across parallel subagents. Furthermore, Opus 5's single \$5/\$25 per million token pricing tier is less flexible than GPT-5.6's three tiers (Sol at \$5/\$30, Terra at \$2.50/\$15, Luna at \$1/\$6), though GPT-5.6 has a billing catch for requests exceeding 272K input tokens.

Key takeaway

For AI engineers evaluating LLMs for daily workflows, especially those involving numerous simple tasks or complex agentic orchestration, you should look beyond overall benchmark scores. Prioritize models that offer specialized execution modes, like GPT-5.6 Sol's "ultra mode" for parallel processing, and flexible, tiered pricing structures. This approach ensures cost-effectiveness and optimal performance for your most frequent operational needs, rather than defaulting to the highest-ranked model.

Key insights

Overall benchmark leadership does not guarantee superiority for specific, common daily workflows.

Principles

Method

GPT-5.6 Sol's "ultra mode" splits tasks across parallel subagents, improving Terminal-Bench 2.1 scores from 88.8% to 91.9%.

In practice

Topics

Best for: AI Architect, NLP Engineer, CTO, AI Engineer, Machine Learning Engineer, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by LLM on Medium.