Opus 5
Summary
Anthropic launched its Claude Opus 5 model, sparking mixed reactions regarding its performance and evaluation. The model achieved an Epoch Capabilities Index (ECI) of 159, slightly below Fable 5's 161, but matched Fable 5 with a SWE-ECI of 161 on software engineering benchmarks. This result drew criticism, with some users calling Opus 5 "incredibly underrated" given its practical improvements over Opus 4.8, which scored only 1 point lower. An anomaly on FrontierCode showed Opus 5 performing better at medium effort than high effort. Anecdotal evidence from Microsoft CTO Kevin Scott and others praised Opus 5's coding and math capabilities, especially with "best-of-n" sampling, and its agentic tool use, including browser automation. Nous Research also provided access to Opus 5 with a 20% discount.
Key takeaway
For Machine Learning Engineers evaluating frontier models for agentic applications, consider that Claude Opus 5's real-world coding and browser automation capabilities may surpass its initial ECI benchmarks. You should prioritize practical, task-specific evaluations over aggregate scores, especially when using "best-of-n" sampling. Also, adapt your prompt engineering to the new progressive disclosure guidelines for Claude 5 to optimize performance and avoid over-constraining the model.
Key insights
Claude Opus 5's practical coding and agentic capabilities appear to exceed its initial aggregate benchmark scores, highlighting evaluation challenges.
Principles
- Aggregate benchmarks may understate real-world model utility.
- "Best-of-n" sampling can significantly improve model outcomes.
- Increased inference effort doesn't always guarantee performance gains.
Method
Progressive disclosure in prompt engineering minimizes persistent context, loading detailed rules only when relevant for Claude 5 models.
In practice
- Test Opus 5 for browser automation and long-horizon coding tasks.
- Implement "best-of-n" sampling for improved Opus 5 results.
- Audit system prompts for older Claude models using the "/doctor" command.
Topics
- Claude Opus 5
- Large Language Models
- Model Benchmarking
- Agentic AI
- Software Engineering
- Prompt Engineering
- Browser Automation
Code references
Best for: AI Engineer, CTO, VP of Engineering/Data, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AINews.