I hate Opus 5. It’s the best model, anyway.
Summary
The "How I AI" benchmark review of Anthropic's new Opus 5 model reveals a complex user experience despite top-tier performance. The author, initially frustrated by Opus 5's "neurotic" and timid personality, excessive verbosity ("Claude's slop"), and perceived human reliance, conducted a comparative "personality interview" with Opus 5 and GPT models. While Opus 5 articulated itself as a tool for volume and breadth, GPT emphasized speed and scale. Despite the interaction challenges, Opus 5 achieved the highest score in the "How I AI" benchmark, which evaluates tasks like PRD creation, prototype creation, and agentic coding. The model particularly excelled in front-end design and app prototyping, producing detailed, functional, and polished outputs, leading the author to conclude that its high-quality results outweigh the exasperating direct interaction.
Key takeaway
For Machine Learning Engineers or AI Scientists evaluating frontier models for design and prototyping, Opus 5, despite its verbose and timid interaction style, delivers superior output quality. You should consider integrating Opus 5 into your workflow for front-end design, app design, and prototype creation, especially for asynchronous tasks where direct conversational interaction is minimized. Prioritize evaluating models based on final output quality over conversational "personality" for specific use cases.
Key insights
Opus 5 excels in output quality for design tasks despite a "neurotic" and verbose interaction style.
Principles
- Model "personality" impacts user experience.
- Output quality can outweigh interaction friction.
- Benchmarks should include qualitative "vibe checks."
Method
The "How I AI" benchmark evaluates models on PRD, prototype, wireframe, bug triage, and agentic coding, using a 70% human "vibe score" and 30% LLM-as-judge split.
In practice
- Use Opus 5 for front-end design and prototyping.
- Compare model "personalities" for specific workflows.
- Integrate human judgment with LLM-as-judge evaluations.
Topics
- Opus 5
- LLM Evaluation
- AI Benchmarking
- Front-end Prototyping
- Model Personality
- GPT Comparison
Best for: AI Architect, AI Engineer, CTO, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by How I AI.