Back to news
launchgoogle2026-09-15

Google DeepMind launches Gemini 3.8 Live and Live Extended Thinking for voice agents

Google DeepMind on September 15 released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, with the Extended Thinking variant topping the Artificial Analysis Speech-to-Speech Index at 82.6.

Google DeepMind 于 9 月 15 日发布两款实时对话模型 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking。两款模型均支持文本、图像、音频与视频输入,可输出文本或音频。

Gemini 3.8 Live 定位为面向大规模、低成本实时对话的可扩展版本,支持近实时视觉输入、97 种语言中自动切换,以及在对话不中断的情况下后台调用工具与 API。Gemini 3.8 Live Extended Thinking 则在对话流中额外叠加一层"后台推理",允许模型在保持语音输出的同时进行多步规划、调用工具并在长任务中向用户口播进度(例如以"Let me check that…"过渡)。

Artificial Analysis 9 月 15 日发布的语音对语音榜显示,Gemini 3.8 Live Extended Thinking 以 82.6 分登顶,超过 OpenAI GPT-Live-1 Astra 的 81.5 与 xAI Grok Voice Think Fast 2.0 High 的 81.3。在复杂任务执行层面,Extended Thinking 在 Tau Voice 基准上达到 68.6%,远超前代 37.7% 的水平。

价格方面,Gemini 3.8 Live 输入音频价格为每分钟 0.005 美元、输出 0.018 美元,折合约 1.38 美元/小时,相比 GPT-Live-1 每小时 3 美元以上的价格低出逾一半。Extended Thinking 变体价格为 3.50 美元/小时,低于 GPT-Live-1 Sol 的 4.47 美元和 Grok Voice Think Fast 2.0 High 的 4.80 美元。

在分发表层,标准版 Gemini 3.8 Live 通过 Gemini API 与 Google AI Studio 对开发者开放,并已在 Search Live 中向普通用户推出;Extended Thinking 版则面向 Google AI 订阅用户在 Gemini App 的 Gemini Live 中提供,并嵌入 Google Workspace 的 Docs、Gmail 与 Keep 语音功能中。

gemini语音extended-thinking实时对话