Deepswe Claims To Measure Agents Better
Summary
DeepSWE, a new software engineering benchmark, has been released to address critical flaws in existing evaluations like SWE-bench Pro. It features 113 original tasks across 91 repositories and five languages, with solutions averaging 668 lines of code. DeepSWE's hand-written verifiers show only 1.4 percent disagreement with LLM judges, significantly improving accuracy over SWE-bench Pro's 8.5 percent false positives and 25 percent false negatives. The benchmark demonstrates 70 percentage points of separation between frontier models, highlighting greater capability gaps than the 30 points seen on SWE-bench Pro. Additionally, DeepSeek permanently reduced its V4 Pro model prices by 75 percent, from \$0.0145–\$3.48 to \$0.003625–\$0.87 per million tokens. Microsoft launched MAI-Image-2.5, a text-to-image model ranking third on the Arena leaderboard, with improvements in text rendering and commercial imagery. Anthropic's Claude Mythos Preview, via Project Glasswing, identified over 10,000 high- or critical-severity vulnerabilities in a month, revealing a significant lag in patching. The Model Context Protocol also proposed major updates, removing stateful sessions and formalizing extensions.
Key takeaway
For AI Engineers evaluating agentic models, consider DeepSWE to gain a more accurate understanding of real-world performance differences. The benchmark's 70 percentage point separation between models suggests existing leaderboards understate capability gaps, impacting deployment decisions. Additionally, security teams should integrate AI-powered vulnerability discovery tools like Claude Mythos Preview, but prepare for a substantial increase in reported issues and prioritize remediation strategies to close the discovery-to-patch window.
Key insights
New benchmarks reveal significant AI agent capability gaps, while AI-driven vulnerability discovery outpaces remediation.
Principles
- Robust benchmarks require original tasks and accurate, purpose-built verifiers.
- AI-powered vulnerability scanning can generate findings faster than human teams can patch.
In practice
- Evaluate agentic models using DeepSWE for more accurate capability assessment.
- Implement AI-assisted vulnerability scanning to identify critical software flaws.
Topics
- AI Benchmarking
- Software Engineering Agents
- DeepSWE
- Vulnerability Discovery
- Large Language Model Pricing
- Text-to-Image Generation
Best for: CTO, VP of Engineering/Data, Research Scientist, AI Engineer, AI Scientist, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai.