Google introduces agentic video understanding in Gemini

Google added agentic video understanding to its latest Gemini models, a capability that moves beyond passive video captioning toward models that can reason over video content and take actions based on it. Google DeepMind said the models 'can now analyze videos with better accuracy,' and the announcement was amplified by CEO Demis Hassabis and the official GoogleDeepMind account.
Technically, agentic video understanding combines temporal reasoning across frames with the tool-use and planning capabilities that define agentic systems — meaning Gemini can not just describe what happens in a video but orchestrate downstream steps based on it. This extends Gemini's already broad multimodal footprint across text, images, and audio into long-form video comprehension, a domain that is computationally demanding and where accuracy has historically been weak.
The feature lands alongside Google's broader consumer AI push, including Google Pics image creation in Workspace and the month's Gemini updates. Competitively, video understanding is a battleground where Google's YouTube-scale data is a natural advantage, and where rivals like OpenAI and Meta (with its Muse voice and perception models) are also investing. Readers should watch for concrete accuracy benchmarks, developer API availability, and how the feature is priced relative to Gemini's aggressive Flash tiers.