Grok just broke the trend

· Source: Matthew Berman · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Software Development & Engineering · Depth: Advanced, quick

Summary

Grok 4.5 has demonstrated exceptional performance on a critical agentic coding benchmark, achieving a score of 83.3. This places it five points ahead of Opus 4.8, indicating a significant advancement in AI-powered coding capabilities. Notably, this impressive result stems from a previous Grok training run, with the next-generation model already in development. This future iteration will leverage extensive cursor data obtained as part of a \$60 billion acquisition. Concurrently, Cursor is advancing its own coding model series, Composer, with Composer 2.5 recognized as a robust workhorse-class model, and Composer 3 anticipated for release soon, further intensifying innovation in this specialized AI domain.

Key takeaway

For AI Engineers evaluating coding assistant models, Grok 4.5's benchmark score of 83.3, surpassing Opus 4.8 by five points, signals a strong contender. You should monitor the upcoming next-generation Grok model, which will integrate acquired cursor data, as it could redefine coding AI capabilities. Also, consider Composer 2.5 for current workhorse coding tasks and anticipate Composer 3's release for enhanced options.

Key insights

Grok 4.5 significantly outperforms competitors on agentic coding benchmarks, with further advancements expected.

In practice

Topics

Best for: Research Scientist, AI Product Manager, Entrepreneur, AI Scientist, Machine Learning Engineer, AI Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Matthew Berman.