Google DeepMind details AGI push with unified Omni model and Gemini Robotics 2

In interviews and research disclosures, Google DeepMind laid out an aggressive AGI roadmap centered on consolidating capabilities into a single multimodal model, Omni, which takes image, audio, video, and text input and generates videos grounded in Gemini's world knowledge. The effort reflects a strategic pivot to treat AI as a core growth driver beyond search, unifying previously separate teams to accelerate the path from research to product.
On the embodied side, Gemini Robotics 2 now controls whole-body movement across a range of machines — from research platforms to Apptronik's Apollo 2 humanoid — extending Gemini from a software brain into a general robot controller. Leadership is also openly discussing self-improving models, notably as Gemini 4 undergoes what DeepMind describes as its most expensive training run to date. Demis Hassabis, fresh off receiving the RSA's Albert Medal, has been pairing the ambition with public reflection on AGI's societal stakes and the DeepMind Institute's cross-disciplinary risk research.
Separately in the voice arena, Google recently shipped Gemini 3.8 Live and an Extended Thinking variant that reasons continuously while speaking across 97 languages, scoring 82.6 on Artificial Analysis' speech-to-speech index. Competitively, DeepMind is positioning Gemini as the one model that spans chat, voice, video generation, and robotics — a breadth play against OpenAI's product sprawl and Anthropic's knowledge-work focus. The open talk of self-improvement, coming the same week as Anthropic's 26% R&D disclosure, cements recursive self-improvement as the theme of the week. Skeptics will note that grand AGI framing from a CEO on an awards circuit is cheap; the Gemini 4 training run's actual results are what matter.