Gemini 3 5 Flash Pairs Smarts With Speed
Summary
Google has launched Gemini 3.5 Flash, an updated mid-tier multimodal model offering significant improvements in agentic capabilities, visual understanding, and speed. This new version, priced three times higher than its predecessor Gemini 3 Flash, supports text, images, audio, and video inputs up to 1 million tokens, and text outputs up to 64,000 tokens at 204 tokens per second. Built on a mixture-of-experts transformer architecture, it features adjustable reasoning levels and "thought preservation" for multi-turn conversations. Gemini 3.5 Flash tops Artificial Analysis's APEX-Agents-AA benchmark with 47.1% accuracy and MMMU-Pro multimodal benchmark with 84% accuracy, though it trails leading models in overall intelligence, knowledge, and coding. It is available free via the Gemini app and Google AI Studio, with API pricing at \$1.50/\$0.15/\$9.00 per million input/cached/output tokens. This release, alongside updates to Antigravity and the introduction of Omni Flash, redefines Google's "Flash" tier as a mid-range offering.
Key takeaway
For AI Engineers developing agentic or low-latency multimodal applications, you should evaluate Gemini 3.5 Flash. Its top performance on APEX-Agents-AA and high speed make it suitable for complex multi-turn tasks and real-time interactions, despite its higher per-token cost compared to previous Flash models. Consider its adjustable reasoning levels and "thought preservation" feature to optimize your agent's performance and context retention.
Key insights
Gemini 3.5 Flash redefines Google's "Flash" tier, offering enhanced agentic and visual capabilities at a higher cost, reflecting a trend of rising token prices.
Principles
- Speed and agentic capability often justify higher token costs.
- Multimodal pretraining enhances diverse task performance.
- Reasoning levels impact benchmark performance significantly.
Method
Gemini 3.5 Flash is a mixture-of-experts transformer, multimodally pretrained on diverse data, then fine-tuned via reinforcement learning for multi-step reasoning and problem-solving.
In practice
- Utilize "thought preservation" for complex multi-turn agentic tasks.
- Adjust reasoning levels to optimize performance for specific use cases.
- Explore Gemini 3.5 Flash for low-latency applications like chatbots.
Topics
- Gemini 3.5 Flash
- Multimodal Models
- Agentic AI
- Mixture-of-Experts
- Large Language Models
- Token Pricing
Best for: NLP Engineer, Computer Vision Engineer, CTO, AI Engineer, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai.