Gemini 3 5 Flash Pairs Smarts With Speed

· Source: The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Advanced, short

Summary

Google has launched Gemini 3.5 Flash, an updated mid-tier multimodal model offering significant improvements in agentic capabilities, visual understanding, and speed. This new version, priced three times higher than its predecessor Gemini 3 Flash, supports text, images, audio, and video inputs up to 1 million tokens, and text outputs up to 64,000 tokens at 204 tokens per second. Built on a mixture-of-experts transformer architecture, it features adjustable reasoning levels and "thought preservation" for multi-turn conversations. Gemini 3.5 Flash tops Artificial Analysis's APEX-Agents-AA benchmark with 47.1% accuracy and MMMU-Pro multimodal benchmark with 84% accuracy, though it trails leading models in overall intelligence, knowledge, and coding. It is available free via the Gemini app and Google AI Studio, with API pricing at \$1.50/\$0.15/\$9.00 per million input/cached/output tokens. This release, alongside updates to Antigravity and the introduction of Omni Flash, redefines Google's "Flash" tier as a mid-range offering.

Key takeaway

For AI Engineers developing agentic or low-latency multimodal applications, you should evaluate Gemini 3.5 Flash. Its top performance on APEX-Agents-AA and high speed make it suitable for complex multi-turn tasks and real-time interactions, despite its higher per-token cost compared to previous Flash models. Consider its adjustable reasoning levels and "thought preservation" feature to optimize your agent's performance and context retention.

Key insights

Gemini 3.5 Flash redefines Google's "Flash" tier, offering enhanced agentic and visual capabilities at a higher cost, reflecting a trend of rising token prices.

Principles

Method

Gemini 3.5 Flash is a mixture-of-experts transformer, multimodally pretrained on diverse data, then fine-tuned via reinforcement learning for multi-step reasoning and problem-solving.

In practice

Topics

Best for: NLP Engineer, Computer Vision Engineer, CTO, AI Engineer, Machine Learning Engineer, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai.