Oh no (the new Grok model is good)
Summary
SpaceX XAI has released Grok 4.5, a new large language model developed in partnership with Cursor, making bold claims about its performance for development work at a significantly reduced cost. Benchmarks from the Artificial Analysis Code Index place Grok 4.5 neck-and-neck with GPT55 and just below Fable, while surpassing Opus 48. Trained on tens of thousands of Nvidia GB300 GPUs with 1.5 trillion parameters, Grok 4.5 excels in coding, agentic tasks, and knowledge work, demonstrating high token efficiency. Its pricing is notably competitive at \$2 per million input tokens and \$6 per million output tokens for contexts under 200,000 tokens, making it 5-10 times cheaper than Fable. Real-world testing confirms its impressive ability to handle complex, multi-step software engineering tasks, including auditing code, generating pull requests, and even creating 3D game environments, often outperforming other models in specific areas like 3D modeling. While it currently lacks the advanced orchestration capabilities seen in models like Fable and GPT56, Grok 4.5 represents a substantial leap for SpaceX XAI, positioning it as a strong competitor in the AI model landscape.
Key takeaway
For AI Engineers or ML teams evaluating new models for software development and agentic workflows, Grok 4.5 presents a compelling, cost-effective option. Its high token efficiency and competitive pricing (starting at \$2/million input tokens) mean you can achieve near-frontier performance for complex coding tasks and 3D content generation at a fraction of the cost of alternatives. Consider integrating Grok 4.5 as your primary code generation and auditing tool, but be aware of its current limitations in advanced multi-agent orchestration.
Key insights
Grok 4.5 offers near-frontier intelligence for coding and agentic tasks at a significantly lower cost and higher token efficiency.
Principles
- Data curation is critical for model quality.
- Token efficiency reduces operational costs.
- Joint training enhances model specialization.
Method
Grok 4.5 was trained on a broad, high-quality dataset spanning STEM and knowledge work, utilizing highly asynchronous RL training on hundreds of thousands of multi-step software engineering tasks with automated and model-based grading.
In practice
- Use Grok 4.5 as a default code model.
- Apply for multi-step software engineering tasks.
- Explore its 3D modeling capabilities.
Topics
- Grok 4.5
- Large Language Models
- Code Generation
- Agentic AI
- Model Benchmarking
- AI Pricing
- 3D Modeling
Best for: CTO, VP of Engineering/Data, Entrepreneur, AI Engineer, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Theo - t3․gg.