Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way
Summary
Google DeepMind has released three new proprietary AI models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, designed for enhanced token efficiency and agentic capabilities. Gemini 3.6 Flash is priced at \$1.50 per million input tokens and \$7.50 per million output tokens, while Gemini 3.5 Flash-Lite offers a lower cost of \$0.30/\$2.50 per million tokens. These models aim to make AI agents faster and cheaper at scale, with 3.6 Flash achieving up to 65% token savings on long-horizon software engineering tasks like DeepSWE, where it scores 49%. Gemini 3.5 Flash-Lite is Google's fastest 3.5 series model, processing 350 output tokens per second. The specialized Gemini 3.5 Flash Cyber is for cybersecurity research, available exclusively to governments and trusted partners. All models are closed-source, API-only, and feature a 1-million-token input context window.
Key takeaway
For AI Engineers and ML Directors optimizing agentic workflows, Google's new Flash models offer compelling cost-performance trade-offs. You should evaluate Gemini 3.6 Flash for complex engineering tasks requiring high token efficiency, or Gemini 3.5 Flash-Lite for high-throughput, low-latency applications. Be aware that these proprietary models entail vendor lock-in and restricted deployment flexibility, especially the specialized 3.5 Flash Cyber, which is only available to trusted partners.
Key insights
Google's new Flash models prioritize token efficiency and speed for AI agents, significantly reducing operational costs and improving performance.
Principles
- Token efficiency directly reduces AI operational costs.
- Specialized models enhance performance for specific domains.
- Proprietary API models restrict deployment flexibility.
In practice
- Use Gemini 3.6 Flash for complex coding and multimodal processing.
- Deploy Gemini 3.5 Flash-Lite for high-throughput, low-latency agentic search.
- Consider Gemini 3.5 Flash Cyber for cybersecurity vulnerability patching.
Topics
- Gemini 3.6 Flash
- AI Agents
- Token Efficiency
- API Pricing
- Cybersecurity AI
- Proprietary Models
Best for: CTO, VP of Engineering/Data, AI Architect, AI Engineer, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by VentureBeat.