Skip to main content

Google Details Gemini 3.5 Audio Models and Native Audio Intelligence

18 SEPTEMBER 2026·2 MIN READ·1 SOURCE·Official source

Google is shifting its translation and transcription systems to process audio directly, launching Gemini 3.5 Live Translate and Gemini 3.5 Transcribe to capture tone, emotion, and conversational code-switching.

Google Details Gemini 3.5 Audio Models and Native Audio Intelligence

Key takeaways · 3

  • 01

    Gemini 3.5 Live Translate supports real-time translation across 70 languages and 2,000 language combinations.

  • 02

    Gemini 3.5 Transcribe powers Gboard's 'Rambler' feature to automatically clean up filler words and correct grammar.

  • 03

    Google trained its Universal Speech Model (USM) on 12 million hours of audio data.

Direct Audio Processing

Google has introduced "native audio intelligence," a technology that processes audio directly rather than relying on a multi-step text conversion pipeline. [1] This approach allows models like Gemini to understand both the sound and the intent simultaneously. [1] Gemini 3.5 Live Translate utilizes this to offer real-time speech translation across 70 languages and over 2,000 language combinations. [1] The system is designed to reflect emotional nuances and natural language transitions like code-switching. [1]

Advanced Transcription

Google's Gemini 3.5 Transcribe operates in environments with heavy background noise and complex terminology to produce formatted text. [1] This model currently powers the "Rambler" feature in Android's Gboard, which removes filler words, corrects grammar, and enables voice-based text editing. [1] As part of its goal to support the world's 1,000 most used languages, Google trained its Universal Speech Model (USM) on 12 million hours of audio data. [1]

What it means

By moving away from traditional speech-to-text-to-speech pipelines, Google is treating audio as a primary modality, allowing models to capture non-verbal cues like tone and emotion. [1] The integration of these direct-processing models into everyday tools like Gboard and real-time translation services demonstrates a focus on practical, multimodal applications. [1] What the sources don't address: Whether the Gemini 3.5 audio features are currently available globally or rolling out in specific regions first.

The shift to 'native audio intelligence' allows models to process tone, pacing, and code-switching directly without losing context in an intermediate text pipeline. This enables more natural voice interfaces and real-time translation.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 18 September 2026

    Google Details Gemini 3.5 Audio Models and Native Audio Intelligence

  2. 18 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.