Google Ships Gemini 3.8 Live and Live Extended Thinking Voice Models

Google shipped Gemini 3.8 Live and 3.8 Live Extended Thinking, native speech-to-speech models built for production-grade real-time voice agents. Available in the Gemini Live API and AI Studio, they're priced at $0.005 per minute of audio input and $0.018 per minute of output — roughly $0.84/hour input, which the community flagged as the cheapest available for scalable voice agents. The 'Extended Thinking' variant adds near-real-time reasoning for complex, multi-step voice tasks.
On benchmarks, Gemini 3.8 Live took the #1 spot on the Speech-to-Speech Index at 82.6 and scored 68.6% on Tau Voice, giving Google a quantified quality lead to pair with its pricing advantage. Native speech-to-speech (rather than a chained STT→LLM→TTS pipeline) is the key architectural bet, cutting latency and preserving prosody and interruption handling that matter for natural conversation.
The launch reframes the voice-agent competitive landscape. OpenAI's realtime API and the newly-shipped Grok Voice on fal now face a Google offering that leads on both a public benchmark and cost — a rare combination. For developers building call-center automation, voice assistants, and conversational agents, $0.84/hour production pricing changes the unit economics.
Google bundled the release with other research announcements — AlphaGenome Atlas mapping all 9 billion single-letter genetic changes in the human genome, and WeatherNext 3 with 50% more accurate precipitation forecasts — underscoring DeepMind's breadth beyond chat. The community praise focused on 'production-readiness,' a notable contrast to earlier voice models judged as demos. The caveat: benchmark leadership on the Speech-to-Speech Index is one axis; real-world robustness across accents, noise, and languages is another. Watch adoption in production voice deployments and whether the pricing holds as usage scales.