Skip to main content

Google Previews Gemini 3.5 Transcribe for Intelligent Real-Time Speech-to-Text

27 AUGUST 2026·2 MIN READ·4 SOURCES·Official source plus independent coverage

Google has unveiled Gemini 3.5 Transcribe, a precise speech-to-text model that actively cleans up filler words and powers advanced voice capabilities.

Google Previews Gemini 3.5 Transcribe for Intelligent Real-Time Speech-to-Text

Key takeaways · 3

  • 01

    Gemini 3.5 Transcribe auto-formats text by removing filler words and handling speaker self-corrections.

  • 02

    The model features two APIs optimized for real-time bidirectional streaming and pre-recorded audio processing.

  • 03

    Function calling capabilities allow the transcription layer to delegate tasks like image generation to other models.

The Launch and Capabilities

Google has officially introduced Gemini 3.5 Transcribe, a new speech-to-text model specifically designed for precise and intelligent real-time transcription. [4] The system converts raw audio directly into auto-formatted text by seamlessly handling speaker self-corrections and removing filler words such as "ums" and "ahs." [1][4] Gemini 3.5 Transcribe already powers the new Rambler capability on Android devices as well as the Gemini application on macOS. [4] Furthermore, the speech-to-text model is coming to the Chrome browser and is currently available in public preview for both developers and enterprises. [3]

Developer Access and APIs

Developers can access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. [4] To support different developer workflows, the model is available across two separate application programming interfaces. [4] A real-time streaming API, specifically named gemini-3.5-transcribe-live, provides continuous, bidirectional streaming with sub-second latency for interactive voice applications. [4] A second Interactions API, called gemini-3.5-transcribe, is distinctly designed to process pre-recorded audio, meetings, and call logs. [4] This pre-recorded audio processing tool includes built-in features for generating speaker attribution and word-level timestamps. [4]

Advanced Voice Functionality

Gemini 3.5 Transcribe features function calling capabilities that allow the transcription model to delegate complex tasks to other Gemini models. [4] According to Google, examples of these delegable tasks include generating images and analyzing files. [4] This specific function calling capability is currently available within the Gemini macOS application. [4] Ultimately, the transcription system is purposefully designed to capture a user's natural speaking style to better understand their intent and recognize custom vocabulary. [4] The tool is designed to convert raw audio directly into accurate, polished text without struggling with complex jargon. [4]

What it means

Google is advancing its speech recognition strategy by baking editing and reasoning capabilities directly into the transcription layer. By offering sub-second latency and function calling, Google positions Gemini 3.5 Transcribe as a powerful endpoint for voice agents and automated workflows. The separation into live and pre-recorded APIs allows developers to optimize for either speed or granular attribution depending on their specific use cases. What the sources don't address: How the model handles complex background noise in real-world enterprise environments compared to conventional speech recognition tools.

By integrating reasoning and editing capabilities directly into the transcription layer, Google is transforming speech-to-text from a passive logging tool into an active agentic interface. This enables cleaner data pipelines and faster task execution for voice-driven enterprise applications.

Why it matters
Story quiz

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 27 August 2026

    Google Previews Gemini 3.5 Transcribe for Intelligent Real-Time Speech-to-Text

  2. 27 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.