What's going on with Gemini?
Summary
Google's AI strategy, particularly concerning Gemini, appears distinct from competitors like Anthropic and OpenAI, despite Google's deep research, custom silicon, and vast resources. While Anthropic and OpenAI lead in frontier model intelligence with models like GPT5.5 and Opus 4.8, Google's Gemini 3.1 Pro lags slightly in benchmarks, and even behind some Chinese models (GLM 5.1, Qwen 3.7) for software engineering tasks. The recent Gemini 3.5 Flash, announced at Google I/O, exhibits underwhelming coding benchmarks but is exceptionally fast, achieving 206 tokens/second, roughly 4x faster than Opus 4.8 and GPT-5.5. However, its price increased significantly to \$9/MTok, making its external market fit unclear. This model likely targets Google's immense internal token consumption, leveraging its custom TPU 8i hardware for inference efficiency. A key weakness for Google is its fragmented coding agent strategy, with tools like Antigravity, Jules, Gemini Code Assist, and AI Studio, which contrasts with unified offerings like Claude Code and Codex, potentially hindering telemetry and training data collection.
Key takeaway
For AI Engineers evaluating large language models for user-facing applications, you should consider Gemini 3.5 Flash's exceptional speed (206 tokens/second) despite its higher cost (\$9/MTok) and mid-pack coding benchmarks. Its design suggests Google optimizes for internal, high-volume use cases leveraging custom hardware. However, if your team relies on integrated coding agents, Google's fragmented tooling ecosystem (Antigravity, AI Studio) presents a significant weakness compared to unified offerings like Claude Code or Codex, potentially impacting your development workflow and data collection.
Key insights
Google's Gemini strategy prioritizes internal token consumption and hardware integration over direct competition in external frontier model intelligence.
Principles
- Hardware-software co-design yields efficiency.
- Internal consumption drives model design.
- Unified developer tools gather data.
In practice
- Evaluate models on speed for user-facing apps.
- Consider self-hosting Chinese models for cost.
- Prioritize integrated developer tooling.
Topics
- Gemini 3.5 Flash
- Large Language Models
- AI Hardware
- Inference Efficiency
- Coding Agents
- Google AI Strategy
Best for: CTO, VP of Engineering/Data, AI Architect, AI Engineer, Director of AI/ML, AI Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Martin Alderson.