Google Launches Gemini 3.5 Transcribe with Smart Disfluency Cleanup
Google has released Gemini 3.5 Transcribe, a new speech-to-text model designed for intelligent voice interactions that automatically removes filler words and formats raw audio into polished text.

Key takeaways · 3
- 01
Gemini 3.5 Transcribe automatically removes "ums" and "ahs" and handles self-corrections in real time.
- 02
The model includes function calling to delegate tasks like image generation to other Gemini models.
- 03
Developers can access the model via real-time streaming and pre-recorded audio processing APIs.
Smart Transcription Capabilities
On August 26, 2026, Google introduced Gemini 3.5 Transcribe, describing it as its most precise speech-to-text model yet for intelligent voice interactions. [4] Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. [4]
The smart transcription feature handles self-corrections, auto-formats text, and edits out filler words like 'ums' and 'ahs'. [1][4] Furthermore, the model can delegate complex tasks, such as image generation and file analysis, to other Gemini models via function calls. [4]
Developer APIs and Integration
The model is currently in public preview for developers and enterprises. [3] Developers can build voice agents, real-time captioning tools, or post-call analytics pipelines with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. [4]
Google made the model available across two separate APIs to suit different developer workflows. [4] The real-time streaming API utilizes `gemini-3.5-transcribe-live` to deliver continuous, bidirectional streaming with sub-second latency for interactive voice apps. [4] Meanwhile, the Interactions API uses `gemini-3.5-transcribe` for pre-recorded audio processing, transcribing meetings and call logs with speaker attribution and word-level timestamps. [4]
Consumer Product Rollout
Consumers are already benefiting from this transcription model with new voice capabilities across products like the Gemini app and on Android. [4] Specifically, the speech-to-text model powers Gboard Rambler and is slated to come to Chrome. [3]
The new voice capabilities are also active via Rambler on Android and in the Gemini app on macOS. [4] According to Google, the function calling capability that delegates tasks to other models is currently available in the Gemini macOS app. [4] The system is designed to capture a user's natural speaking style to better understand intent and recognize custom vocabulary, enabling users to execute tasks with their voice. [4]
What it means
By integrating disfluency cleanup and function calling natively into the transcription model, Google is moving beyond conventional speech recognition systems that struggle with jargon and background noise. Packaging these features into a streaming API for sub-second latency suggests a strong push to enable interactive voice agents. The rollout spans multiple consumer entry points, including Gboard Rambler, Android, macOS, and Chrome, establishing a wide deployment footprint for the new architecture. What the sources don't address: How the operational costs of running these advanced transcription models will impact pricing tiers for developers using the API.
The integration of formatting and function calling natively into a speech-to-text model reduces the need for downstream processing pipelines. This allows developers to build more responsive and accurate voice interfaces with lower latency.
Why it matters
Test yourself on this story — 5 questions.
Create a free account to take the quiz, earn XP, and get a daily session built for your industry.
Take the quizHow this developed
3 September 2026
Archived
27 August 2026
Google Previews Gemini 3.5 Transcribe for Intelligent Real-Time Speech-to-Text
27 August 2026
Event created from source cluster.
Sources
- Introducing Gemini 3.5 Transcribeblog.google
- Google’s new AI transcription edits out your ‘ums’ and ‘ahs’The Verge AI
- Google announces Gemini 3.5 Transcribe for AI-powered speech-to-textArs Technica AI
- Google debuts Gemini 3.5 Transcribe, a speech-to-text model that powers Gboard Rambler and is coming to Chrome, in public preview for developers and enterprises (Abner Li/9to5Google)Techmeme