Google’s Gemini 3.8 Live Brings Deeper Reasoning to Real-Time Voice AI

 

Image: Google

Gemini 3.8 Live focuses on fast, natural conversations, while a new Extended Thinking model is designed to keep talking as it works through more complicated, multi-step tasks.

Google is expanding Gemini’s voice capabilities with Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two native audio models built to make real-time AI conversations more capable without turning every complex request into a long silent pause.

Announced on September 15, 2026, the models take two different approaches to live interaction. Gemini 3.8 Live is Google’s lower-latency, cost-focused option for everyday voice agents, while Gemini 3.8 Live Extended Thinking adds more substantial background reasoning for tasks that involve planning, multiple steps or external tools.

The more consequential change isn’t simply better speech generation. Google is trying to make voice agents capable of doing work while the conversation continues.

Two models for two kinds of voice interaction

Gemini 3.8 Live is designed as the default choice for low-latency voice applications. It can accept text, images, audio and video, and respond using text or generated audio. Google says the model supports visual grounding, meaning a conversation can incorporate what the system is seeing through a camera or other visual input rather than relying on voice alone.

The model can also initiate function and API calls asynchronously. Instead of freezing a conversation whenever an external service needs to return information, the agent can continue interacting with the user while that operation runs in the background.

That could matter for applications such as customer service, travel planning or technical support. A voice agent might begin checking a reservation, for example, while continuing to clarify what the user wants instead of forcing the conversation to stop until the lookup finishes.

Gemini 3.8 Live Extended Thinking takes the same idea further. Google designed it for requests requiring additional reasoning, planning and multi-step execution. The model can perform background reasoning and asynchronous tool calls while continuing to generate spoken responses.

In practice, Google says Extended Thinking can provide short spoken progress updates as it works through a longer task. That makes the interaction closer to an ongoing conversation than the familiar pattern of asking an AI a difficult question and waiting silently for a finished answer.

Gemini 3.8 Live vs. Extended Thinking

Gemini 3.8 LiveGemini 3.8 Live Extended Thinking
Primary focusFast, scalable voice interactionComplex, multi-step voice tasks
Background reasoningInterleaved reasoningDeeper configurable reasoning
Function callingAsynchronous supportedAsynchronous
Visual inputYesYes
Input limit131,072 tokens131,072 tokens
Maximum output65,536 tokens65,536 tokens
Best suited forGeneral voice agents, support, live assistancePlanning, agentic workflows, complex tool use

Google’s developer documentation lists the same 131,072-token input limit and 65,536-token output limit for both models.

The distinction, then, is less about how much information the models can accept and more about how they allocate computation during a live interaction.

It can keep using tools while you talk

Asynchronous tool use is one of the most practical additions in the new models.

Traditional voice assistants commonly follow a sequential process: listen to a request, stop, call a service, wait for a response and then start speaking again. That works for simple queries, but it becomes awkward when an AI agent needs to coordinate several actions.

Gemini 3.8 Live can instead launch an external function while continuing the audio session. Google also supports incremental content updates, allowing new structured information to be incorporated into an ongoing conversation as it becomes available.

Extended Thinking builds on that system by performing more complex reasoning in parallel with those background operations.

Google demonstrated the model handling multi-step bookings, working from voice feedback and rough sketches to produce React components, and coordinating tasks involving multiple asynchronous function calls. Those demonstrations are controlled examples rather than independent evaluations, but they illustrate the kind of voice-agent workflow Google is targeting.

Gemini can also watch what is happening

The Live models aren't limited to speech.

Gemini 3.8 Live can process visual information in near real time, allowing conversations to reference what a user is showing the model. Google demonstrated this capability in scenarios including employee onboarding and visually guided tasks.

That combination of audio, vision and tools opens up possibilities beyond conventional voice assistants. A support agent could potentially look at a device while helping troubleshoot it, for example, rather than relying entirely on the user's verbal description.

Google also says Gemini 3.8 Live can automatically recognize and transition between 97 supported languages during a conversation, including cases where a speaker switches languages mid-session.

Extended Thinking posts strong early benchmark results

The more reasoning-heavy model also arrives with strong benchmark numbers.

Artificial Analysis currently gives Gemini 3.8 Live Extended Thinking (High) an overall Speech-to-Speech Index score of 82.6. The same independent benchmark reports 98% on its speech-reasoning test and 68.6% on τ-Voice, which measures task completion in simulated customer-service scenarios involving tools.

The standard Gemini 3.8 Live takes a different balance. Artificial Analysis lists it at 76.0 on its overall speech-to-speech index, with lower agentic performance but stronger conversational-dynamics results and faster time to first audio than Extended Thinking in its testing.

That difference reinforces the purpose of Google's two-model strategy: the base version prioritizes conversational responsiveness and cost, while Extended Thinking spends more computation on difficult tasks.

As with any AI benchmark, however, these results describe particular test environments rather than every real-world conversation. Performance can vary substantially depending on prompts, tools, latency, languages and the application surrounding the model.

Google is also pushing down voice-agent costs

Pricing could be another significant part of the release.

Google says developers using its Live API are charged $0.005 per minute for audio input and $0.018 per minute for audio output for the new models.

Artificial Analysis' broader cost benchmark, which incorporates audio, text, outputs and exposed reasoning costs during its test workload, calculated roughly $0.84 per hour of input audio workload for Gemini 3.8 Live and $3.50 for Extended Thinking at its High setting. Those figures are benchmark-specific rather than Google's simple API list prices, so they should not be treated as direct estimates of every application's bill.

For companies operating large numbers of automated voice conversations, however, cost per minute can become as important as raw model capability.

Gemini 3.8 Live is spreading across Google products

These models aren't restricted to developers.

Google is rolling Gemini 3.8 Live into Search Live, while Extended Thinking is being used across parts of the Gemini app and Google Workspace, including voice-based experiences in Gmail, Docs and Keep. Enterprise versions are also being offered through Google's enterprise platforms, with some features beginning in private preview.

Developers can access both models through the Gemini API and Google AI Studio, while Google Cloud and Vertex AI are also listed among their distribution channels.

Google is additionally working with voice and agent infrastructure platforms including LiveKit, LangChain, Agora, Pipecat, Fishjam and Vercel to make the models easier to integrate into third-party applications.

Generated audio gets SynthID watermarking

Google says audio generated by its AI products is embedded with SynthID, the company's imperceptible watermarking technology for identifying AI-generated media.

That matters more as live models become capable of producing increasingly natural speech. Voice interfaces can make AI interaction more accessible, but realistic synthetic audio also creates risks around impersonation and misleading content.

Watermarking isn't a complete solution to those problems, but Google is making it part of the infrastructure surrounding its audio models rather than treating generated speech as an unmarked output.

There are still limitations

Despite the emphasis on more capable reasoning, Google’s own model card notes that Gemini 3.8 Live and Extended Thinking retain familiar limitations of large AI models, including the potential to hallucinate incorrect information. Google also warns that users may encounter occasional slowness or timeouts.

The model card lists a January 2025 knowledge cutoff, meaning applications requiring current information still need appropriate grounding or external tools. Both models support Google Search grounding through the API, according to Google's documentation.

Those constraints are particularly important for agents that can perform actions rather than merely answer questions. A convincing spoken response is not necessarily a correct one, and giving an AI access to external functions raises the stakes when reasoning fails.

Voice AI is becoming less about talking and more about doing

The larger shift behind Gemini 3.8 Live is that Google is treating voice less as an output format and more as an interface for AI agents.

Fast speech generation has already made conversations with AI feel less mechanical. The harder problem is what happens when a spoken request requires research, external services, planning or a sequence of actions.

Gemini 3.8 Live attempts to keep those interactions responsive. Extended Thinking goes further by allowing more reasoning to happen in the background while the spoken conversation continues.

That doesn't eliminate the reliability problems surrounding AI agents, nor does a strong benchmark result guarantee dependable performance in every deployment. But it does move Google's voice models toward a different kind of assistant: one designed not simply to answer while you talk, but to continue the conversation while it works on something else.