Google DeepMind on September 15 launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of speech-to-speech models it says are the first Gemini voice systems capable of reasoning while they talk, and turned them on across Gemini Live, Gmail, Google Docs, Keep and the Search Live experience inside AI Mode.
Reasoning while it speaks
Gemini 3.8 Live is Google's low-latency conversational agent, with visual grounding and asynchronous tool use. The Extended Thinking variant is designed to pause and reason mid-conversation, useful for agentic voice tasks like handling a bank support call, walking a user through a Gmail thread or drafting a Google Doc out loud. On Artificial Analysis's Speech-to-Speech Quality Index the models take the #1 spot with a score of 82.6, and Google reports 68.6% on the τ-Voice agentic benchmark and 97.7% on Big Bench Audio reasoning.
Where it's shipping
Users see Extended Thinking rolling out today inside the Gemini Live app, with conversational search inside Gmail, note-taking in Keep and "draft-out-loud" editing in Google Docs. The base 3.8 Live model powers the Search Live experience inside AI Mode, and both variants are available in the Gemini API and Google AI Studio for developers.
Why this launch matters
Voice agents that can hold up under real-world tool use have become the next battlefield in the model wars — OpenAI ships Advanced Voice, and every major agent company is racing to hide latency behind reasoning. By moving Extended Thinking into a live speech loop, Google is trying to close the gap where audio agents fall over: complex conversations that need memory, tool calls and follow-up questions. The launch also lands alongside the Fairwind cyber variant of Gemini 3.8 Flash and Google's push to fold Gemini into every Workspace surface it owns, from Docs to Google Home devices.
Reporting based on coverage from 9to5Google, Google DeepMind and Artificial Analysis.
