Google ships three new Gemini Flash models but delays frontier Gemini 3.5 Pro

Google DeepMind launched three new Flash-tier models aimed at developers building agents at scale. Gemini 3.6 Flash is positioned as the workhorse for coding and document analysis and uses about 17% fewer output tokens than its predecessor, with Jeff Dean posting side-by-side efficiency demos; 3.5 Flash-Lite targets low-latency tasks; and 3.5 Flash Cyber is a specialized model to detect, validate, and patch software vulnerabilities, offered to governments and trusted partners via CodeMender. But the headline frontier model, Gemini 3.5 Pro, slipped its July 17 target after falling short on coding and reasoning benchmarks — no official model card, pricing, or benchmark scores were released, and Alphabet shares fell roughly 4%.
The technical story is a bifurcation: Google is shipping cheap, fast, efficient models on schedule while its flagship stalls. Gizmodo reported that 3.6 Flash trails Claude Sonnet 5, GPT-5.6, and even xAI's Grok 4.5 on agentic coding. Google is also developing an efficiency chip dubbed 'Frozen v2' (slated for 2028), reportedly 6–10x more efficient than current silicon.
Competitively the delay dents developer trust — Bind AI warned 'do not rebuild architecture around a July date — Google has missed two major delivery targets this year,' and r/MachineLearning advised treating Pro GA as a 'bonus event, not dependency.' Yet Logan Kilpatrick teased that Google has started its 'most ambitious pre-training run yet, for Gemini 4,' and Sundar Pichai touted the Gemini app at 950M MAU and APIs processing 22B tokens/min. Watch whether Pro ships before Gemini 4 overtakes it and whether Flash Cyber sees real security adoption.