Skip to main content

Google Launches Gemini 3.5 Transcribe with Smart Disfluency Cleanup

27 AUGUST 2026·2 MIN READ·4 SOURCES·Official source plus independent coverage

Google has released Gemini 3.5 Transcribe, a new speech-to-text model designed for intelligent voice interactions that automatically removes filler words and formats raw audio into polished text.

Google Launches Gemini 3.5 Transcribe with Smart Disfluency Cleanup

Key takeaways · 3

  • 01

    Gemini 3.5 Transcribe automatically removes "ums" and "ahs" and handles self-corrections in real time.

  • 02

    The model includes function calling to delegate tasks like image generation to other Gemini models.

  • 03

    Developers can access the model via real-time streaming and pre-recorded audio processing APIs.

Smart Transcription Capabilities

On August 26, 2026, Google introduced Gemini 3.5 Transcribe, describing it as its most precise speech-to-text model yet for intelligent voice interactions. [4] Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. [4]

The smart transcription feature handles self-corrections, auto-formats text, and edits out filler words like 'ums' and 'ahs'. [1][4] Furthermore, the model can delegate complex tasks, such as image generation and file analysis, to other Gemini models via function calls. [4]

Developer APIs and Integration

The model is currently in public preview for developers and enterprises. [3] Developers can build voice agents, real-time captioning tools, or post-call analytics pipelines with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. [4]

Google made the model available across two separate APIs to suit different developer workflows. [4] The real-time streaming API utilizes `gemini-3.5-transcribe-live` to deliver continuous, bidirectional streaming with sub-second latency for interactive voice apps. [4] Meanwhile, the Interactions API uses `gemini-3.5-transcribe` for pre-recorded audio processing, transcribing meetings and call logs with speaker attribution and word-level timestamps. [4]

Consumer Product Rollout

Consumers are already benefiting from this transcription model with new voice capabilities across products like the Gemini app and on Android. [4] Specifically, the speech-to-text model powers Gboard Rambler and is slated to come to Chrome. [3]

The new voice capabilities are also active via Rambler on Android and in the Gemini app on macOS. [4] According to Google, the function calling capability that delegates tasks to other models is currently available in the Gemini macOS app. [4] The system is designed to capture a user's natural speaking style to better understand intent and recognize custom vocabulary, enabling users to execute tasks with their voice. [4]

What it means

By integrating disfluency cleanup and function calling natively into the transcription model, Google is moving beyond conventional speech recognition systems that struggle with jargon and background noise. Packaging these features into a streaming API for sub-second latency suggests a strong push to enable interactive voice agents. The rollout spans multiple consumer entry points, including Gboard Rambler, Android, macOS, and Chrome, establishing a wide deployment footprint for the new architecture. What the sources don't address: How the operational costs of running these advanced transcription models will impact pricing tiers for developers using the API.

The integration of formatting and function calling natively into a speech-to-text model reduces the need for downstream processing pipelines. This allows developers to build more responsive and accurate voice interfaces with lower latency.

Why it matters
Story quiz

Test yourself on this story — 5 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 3 September 2026

    Archived

  2. 27 August 2026

    Google Previews Gemini 3.5 Transcribe for Intelligent Real-Time Speech-to-Text

  3. 27 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.