Google launches Gemini 3.8 Live models with background reasoning for voice agents
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new near real-time reasoning models for voice agents. [1][3] The models feature asynchronous function calling and multi-step reasoning capabilities. [1][2]

Key takeaways · 3
- 01
Gemini 3.8 Live enables asynchronous function calling while streaming audio.
- 02
Gemini 3.8 Live Extended Thinking captured the top spot on the Artificial Analysis Speech to Speech Quality Index.
- 03
The models are priced at $0.005/min for audio input and $0.018/min for audio output.
Upgraded voice agents
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking to enable more intelligent near real-time voice agents. [1][3] The 3.8 Live model combines conversational intelligence with visual grounding, while the Extended Thinking model is designed for high-complexity tasks requiring multi-step reasoning. [1][3]
Developers can use the models via the Live API to build agents capable of executing API and tool calls in the background while continuing to stream audio to the user. [2] The Extended Thinking model supports configurable thinking to handle reasoning in the background while responding or narrating progress. [2]
Benchmarks and pricing
Gemini 3.8 Live Extended Thinking captured the number one spot on the Artificial Analysis Speech to Speech Quality Index with a score of 82.6. [1][3] It also led in agentic task completion with 68.6% on the τ-Voice benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark. [1][3]
The models are available via the Live API, with pricing at $0.005 per minute for audio input and $0.018 per minute for audio output. [2] The release expands Google's developer suite for building real-time, voice-first product experiences. [2]
What it means
The introduction of asynchronous function calling in Gemini 3.8 Live addresses a key friction point in voice AI by allowing agents to fetch data without pausing the conversation. The Extended Thinking model's performance on the Artificial Analysis Speech-to-Speech Quality Index positions it strongly against other frontier voice models. What the sources don't address: how the latency of the extended reasoning process affects the user experience in practical voice applications.
The ability to run tool calls asynchronously while streaming audio dialogue represents a significant workflow improvement for voice agent developers. The inclusion of multi-step reasoning allows for more complex, agentic task completion via voice.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
16 September 2026
Event evidence refreshed from source cluster.
15 September 2026
Event evidence refreshed from source cluster.
15 September 2026
Event evidence refreshed from source cluster.
15 September 2026
Google DeepMind Launches Gemini 3.8 Live Voice Models
15 September 2026
Event created from source cluster.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended ThinkingGoogle DeepMind Blog
- New Gemini Audio models for developersblog.google
- Gemini 3.8 Live & Gemini 3.8 Live Extended Thinkingblog.google
- Gemini 3.8 Live Extended Thinking Debuts With Advanced Voice - NewsBricksnewsbricks.com