Google Gemini gains agentic video understanding, up to 88% fewer tokens on long videos
Google launched agentic video understanding on September 1 for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, with up to 88% fewer tokens and 66% lower cost on long-form videos.
Google launched agentic video understanding on September 1 for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The capability replaces fixed-rate frame sampling with an agentic loop that lets the model decide which moments to watch, at what speed, and through which modality: frames, audio, or transcript.
Google's own benchmarks show up to 88% fewer tokens consumed, up to 66% lower cost, and up to 7 percentage points higher accuracy on standard video analysis tests. The efficiency gain grows sharply as video length increases, from 10-minute how-to guides through 90-plus-minute recordings. Pricing follows standard Gemini token rates, with no new feature fee. The capability supports both direct video uploads and YouTube videos.
Customer Ponder, which had built its own agentic video pipeline on Gemini, reported that Google's built-in version matched it at roughly one-third of the input tokens. The capability is available to developers using the Interactions and GenerateContent APIs.