Google unveils Gemini 3.8 Live and Extended Thinking
The company said the models cut voice latency to zero and run tool calls plus multi-step reasoning in parallel while a user speaks.
Mateo Fernandez ·
The company said on October 3, 2026 it released Gemini 3.8 Live and a companion Extended Thinking model, describing them as voice-capable systems that eliminate perceptible latency while executing background tool calls and complex reasoning. Market reaction: Reaction pending.
Low-latency architecture and tool calls
The company said Gemini 3.8 Live maintains an uninterrupted conversational stream as separate processes call external tools and run multi-step chains of thought in parallel. Officials said the architecture routes user audio to a front-end low-latency engine while heavier reasoning and tool invocations run asynchronously, with results merged into the live reply.
That split-process design, the company said, aims to let assistants respond immediately to users while completing longer-running tasks such as database queries, code execution or multi-stage planning behind the scenes. Officials cautioned that the approach raises integration and resource questions: connecting bespoke tools will require API work and the parallel workloads may increase compute and latency variability for partners.
Developers and integrators should expect to test the models against latency and cost benchmarks in the coming days; the company listed October 3, 2026 as the release date. Watch for platform and SDK notes and early performance reports by October 10, 2026.