Oh no (the new Grok model is good)

· Source: Theo - t3․gg · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Software Development & Engineering · Depth: Advanced, extended

Summary

SpaceX XAI has released Grok 4.5, a new large language model developed in partnership with Cursor, making bold claims about its performance for development work at a significantly reduced cost. Benchmarks from the Artificial Analysis Code Index place Grok 4.5 neck-and-neck with GPT55 and just below Fable, while surpassing Opus 48. Trained on tens of thousands of Nvidia GB300 GPUs with 1.5 trillion parameters, Grok 4.5 excels in coding, agentic tasks, and knowledge work, demonstrating high token efficiency. Its pricing is notably competitive at \$2 per million input tokens and \$6 per million output tokens for contexts under 200,000 tokens, making it 5-10 times cheaper than Fable. Real-world testing confirms its impressive ability to handle complex, multi-step software engineering tasks, including auditing code, generating pull requests, and even creating 3D game environments, often outperforming other models in specific areas like 3D modeling. While it currently lacks the advanced orchestration capabilities seen in models like Fable and GPT56, Grok 4.5 represents a substantial leap for SpaceX XAI, positioning it as a strong competitor in the AI model landscape.

Key takeaway

For AI Engineers or ML teams evaluating new models for software development and agentic workflows, Grok 4.5 presents a compelling, cost-effective option. Its high token efficiency and competitive pricing (starting at \$2/million input tokens) mean you can achieve near-frontier performance for complex coding tasks and 3D content generation at a fraction of the cost of alternatives. Consider integrating Grok 4.5 as your primary code generation and auditing tool, but be aware of its current limitations in advanced multi-agent orchestration.

Key insights

Grok 4.5 offers near-frontier intelligence for coding and agentic tasks at a significantly lower cost and higher token efficiency.

Principles

Method

Grok 4.5 was trained on a broad, high-quality dataset spanning STEM and knowledge work, utilizing highly asynchronous RL training on hundreds of thousands of multi-step software engineering tasks with automated and model-based grading.

In practice

Topics

Best for: CTO, VP of Engineering/Data, Entrepreneur, AI Engineer, Machine Learning Engineer, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Theo - t3․gg.