Google Ships Gemini 3.8 Live Voice Models, Topping Speech-to-Speech Index at 82.6

Gemini 3.8 Live is a genuine architectural step beyond chatbots that read text aloud: these are native audio-to-audio models that keep a conversation flowing while running tools in the background and reasoning over live visual input from a camera. The Extended Thinking variant scored 82.6 on Artificial Analysis' Speech-to-Speech Index — the new #1 — while the standard Live model posted 76.0, and both support automatic switching across 97 languages mid-sentence.
The pricing structure is deliberately tiered to make real-time voice economically viable at scale: standard dialogue runs about $0.84/hour (roughly $0.005/min) while complex multi-step reasoning costs $3.50/hour. Developers on Hacker News (319 points, 195 comments) contrasted this favorably against OpenAI's real-time stack, which was compared unfavorably on cost. Availability is via the Gemini API and Google AI Studio.
Google bundled several adjacent releases the same day — a standalone Gemini app for Windows 10/11 with Alt+Space integration, Gemini Notebook education features in nearly 100 languages, and Gemini Robotics 2 for whole-body robot control — but the Live voice models are the headline. The launch drew the top slot on Hacker News, a strong developer signal.
The competitive framing is direct: this is Google pressing its multimodal and cost advantages against OpenAI's advanced voice mode and against the wave of voice-agent startups. The open question is latency-under-load and whether background tool execution stays reliable in production; early praise centered on the demos and pricing rather than sustained real-world deployment. Watch for third-party latency benchmarks over the coming weeks.