Google developing custom 'Frozen v2' AI chip for 2028 to run Gemini 6–10x more efficiently

Google is developing a new custom AI chip, internally dubbed 'Frozen v2,' designed specifically to improve the efficiency and performance of Gemini models, according to reporting from The Information relayed by TechCrunch and Gigazine. The chip is slated for a 2028 release and could deliver 6 to 10 times better efficiency than Google's current chips when measured by tokens generated per unit of power. Google declined to directly confirm the report.
The motivation is straightforward economics: inference costs dominate at Gemini's scale, and power efficiency directly determines margin. A 6–10x gain in tokens-per-watt would dramatically lower the cost of serving Gemini across Search, Workspace, and the API, and reduce Google's dependence on NVIDIA GPUs. It extends Google's long-running TPU strategy — the company has built custom accelerators for years — into a new generation targeted explicitly at the Gemini era.
The move fits a clear theme-of-the-week: everyone is racing to escape NVIDIA dependence. DeepSeek is developing its own inference chip, Apple is reportedly hunting AI-chip acquisitions, and NVIDIA itself is doubling down on NVLink and AI-factory scale-up networking. Custom silicon is becoming the primary lever for cost advantage as model quality commoditizes and Chinese open-weight models compress per-token pricing.
Caveats abound: 2028 is far off, efficiency claims are pre-silicon projections, and 'tokens per unit of power' is a metric Google chose to emphasize. Actual competitive impact depends on execution, fab capacity, and whether NVIDIA's roadmap (Rubin and beyond) closes the gap. Readers should watch for confirmation, benchmark disclosures, and whether Google offers Frozen v2 capacity externally via Google Cloud the way it does current TPUs.