Google has introduced agentic video understanding in Gemini, significantly improving how the AI analyzes long videos while making the process more affordable. The company announced the upgrade for its latest models, including Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, marking another step in Google’s steady effort to make Gemini more useful.
The new capability lets users upload videos and have Gemini analyze them with far greater efficiency than before. Crucially, it reduces token usage by up to 88% while simultaneously offering up to 7% better accuracy. Because these improvements target both cost and performance together, the feature addresses two longstanding pain points for developers working with video content.
Previously, users could upload videos to Gemini, but the AI could only perform what Google calls “static” processing. This meant it would split each video into individual frames, resulting in slower performance and considerably higher costs. Since every frame required separate analysis, longer videos became expensive and time-consuming to process effectively.
With agentic video understanding, Gemini can now decide what to watch and at what speed. Additionally, it can choose between frames, audio, and the transcript to analyze content more intelligently. This means Gemini can pinpoint split-second changes, answer complex questions across multi-hour videos, and inspect footage for visual artifacts. The model can even count and track physical movements and distinct objects within videos.
For now, agentic video understanding remains available only through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. However, Google confirmed the feature will roll out to the Gemini app soon, expanding access beyond developers and enterprise users. Because the capability handles multi-hour videos without breaking a sweat, its arrival on consumer platforms could meaningfully broaden how people interact with video content.
Google also plans to use agentic video understanding to power YouTube’s “Ask YouTube” feature going forward. This integration could make it substantially easier for creators to analyze their own videos directly through Gemini. As a result, the upgrade positions Gemini as a stronger tool for both technical developers building video applications and everyday creators seeking deeper insights from their content across Google’s expanding ecosystem of AI-powered services.
