Skip to main content

Google Introduces Gemini 3.8 Live and Extended Thinking Audio Models

17 SEPTEMBER 2026·3 MIN READ·10 SOURCES·Official source plus independent coverage

Google has launched two new conversational AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed for near real-time voice interaction and multi-step reasoning.

Google Introduces Gemini 3.8 Live and Extended Thinking Audio Models

Key takeaways · 4

  • 01

    Gemini 3.8 Live Extended Thinking can perform background API calls while continuing verbal dialogue.

  • 02

    The models support up to 128K input tokens across audio, image, video, and text.

  • 03

    Audio pricing is set at $0.005 per minute for input and $0.018 per minute for output.

  • 04

    All generated audio includes Google DeepMind's SynthID invisible watermarking.

New Conversational Models

Google announced the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced live dialogue models. [4] Both systems are built on the Gemini 3 Pro base model and feature a maximum input context window of 128K tokens that can ingest audio, images, video, and text. [8] The models generate outputs in both audio and text formats, supporting a maximum output window of 64K tokens. [8]

Gemini 3.8 Live is optimized for scale, cost efficiency, and low-latency voice interactions. [1][7] The system is capable of visual grounding, allowing it to process visual inputs from a camera in near real time to provide context for conversations. [1] Additionally, the model can automatically detect and switch between 97 supported languages during a single conversation without requiring the user to manually adjust settings. [3][5]

Extended Thinking Capabilities

Gemini 3.8 Live Extended Thinking is designed to handle complex tasks by performing multi-step reasoning simultaneously with speech generation. [1] The model can execute asynchronous tool usage and background API calls while maintaining an active conversation, emitting natural verbal cues like "Let me check that" to indicate it is processing a request. [2][7] This approach eliminates long pauses while the system retrieves data from multiple sources or completes multi-step workflows. [7]

In industry benchmarks, the Extended Thinking model achieved a score of 82.6 on the Artificial Analysis Speech to Speech Quality Index, placing it first overall in that metric. [1][2] The model also recorded a 97.7% score on the Big Bench Audio benchmark and a 68.6% completion rate on the T-Voice test. [2] However, the model achieved a 35.1% completion rate on the Sierra τ-Voice-banking benchmark, indicating that performance in complex financial scenarios still has room for improvement. [1][2]

Pricing and Ecosystem Integration

The new models are currently available to developers through the Gemini API and Google AI Studio, and to enterprise customers via private preview in Gemini Enterprise. [1][2] Google set the pricing for the models at $0.005 per minute for audio input and $0.018 per minute for audio output, which roughly equates to $3 per million input tokens and $12 per million output tokens. [8] Developers can access the models through integration partners including Agora, LiveKit, Vercel, Pipecat, Fishjam, and Vision Agents. [2]

The models are being integrated into Google Workspace applications, enabling users to retrieve information and edit documents using voice commands in Gmail, Google Docs, and Google Keep. [6] To address risks related to voice cloning and misinformation, all audio generated by Gemini 3.8 Live and Extended Thinking includes Google DeepMind's invisible SynthID watermark, which allows detection tools to identify AI-generated content. [1][2]

Expanded Audio Portfolio

Alongside the live dialogue models, Google highlighted several other updates to its portfolio of voice and audio systems. [8] The company detailed Gemini 3.5 Transcribe, a speech recognition model supporting over 85 languages with an average word error rate (WER) of 4.0% for streaming and 2.6% for non-streaming applications. [8]

Google also detailed the Gemini 3.5 Live Translate model, which supports voice translation across more than 70 languages. [8] The company's audio generation lineup further includes the Gemini 3.1 Flash TTS model for text-to-speech synthesis and the Lyria 3.5 model designed for music generation. [8]

What it means

Google's rollout of the Gemini 3.8 Live series represents a structural shift away from cascaded voice systems—where speech recognition, text generation, and text-to-speech operate sequentially—toward parallel processing. By enabling the Extended Thinking model to handle background tool execution while producing natural conversational filler, Google aims to eliminate the latency that typically plagues complex agentic workflows. The 128K multimodal context window and visual grounding directly challenge the capabilities of competing systems like OpenAI's GPT-4o Realtime API and Amazon's Nova Sonic. What the sources don't address: How the model's energy consumption and inference latency scale when simultaneously processing maximum-length visual context alongside multi-step asynchronous tool use.

The transition to parallel reasoning and speech generation allows voice agents to handle long-running tasks without awkward conversational pauses. This enables developers to build more fluid, real-time voice applications that can interact with APIs and visual data simultaneously.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 17 September 2026

    Google Introduces Gemini 3.8 Live and Extended Thinking Audio Models

  2. 17 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.