Skip to main content

Google launches Gemini 3.8 Live models with background reasoning for voice agents

15 SEPTEMBER 2026·2 MIN READ·4 SOURCES·Official source plus independent coverage

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new near real-time reasoning models for voice agents. [1][3] The models feature asynchronous function calling and multi-step reasoning capabilities. [1][2]

Google launches Gemini 3.8 Live models with background reasoning for voice agents

Key takeaways · 3

  • 01

    Gemini 3.8 Live enables asynchronous function calling while streaming audio.

  • 02

    Gemini 3.8 Live Extended Thinking captured the top spot on the Artificial Analysis Speech to Speech Quality Index.

  • 03

    The models are priced at $0.005/min for audio input and $0.018/min for audio output.

Upgraded voice agents

Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking to enable more intelligent near real-time voice agents. [1][3] The 3.8 Live model combines conversational intelligence with visual grounding, while the Extended Thinking model is designed for high-complexity tasks requiring multi-step reasoning. [1][3]

Developers can use the models via the Live API to build agents capable of executing API and tool calls in the background while continuing to stream audio to the user. [2] The Extended Thinking model supports configurable thinking to handle reasoning in the background while responding or narrating progress. [2]

Benchmarks and pricing

Gemini 3.8 Live Extended Thinking captured the number one spot on the Artificial Analysis Speech to Speech Quality Index with a score of 82.6. [1][3] It also led in agentic task completion with 68.6% on the τ-Voice benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark. [1][3]

The models are available via the Live API, with pricing at $0.005 per minute for audio input and $0.018 per minute for audio output. [2] The release expands Google's developer suite for building real-time, voice-first product experiences. [2]

What it means

The introduction of asynchronous function calling in Gemini 3.8 Live addresses a key friction point in voice AI by allowing agents to fetch data without pausing the conversation. The Extended Thinking model's performance on the Artificial Analysis Speech-to-Speech Quality Index positions it strongly against other frontier voice models. What the sources don't address: how the latency of the extended reasoning process affects the user experience in practical voice applications.

The ability to run tool calls asynchronously while streaming audio dialogue represents a significant workflow improvement for voice agent developers. The inclusion of multi-step reasoning allows for more complex, agentic task completion via voice.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 16 September 2026

    Event evidence refreshed from source cluster.

  2. 15 September 2026

    Event evidence refreshed from source cluster.

  3. 15 September 2026

    Event evidence refreshed from source cluster.

  4. 15 September 2026

    Google DeepMind Launches Gemini 3.8 Live Voice Models

  5. 15 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.