Skip to main content

Google Releases Gemini 3.5 Transcribe with 2.6% WER and Dual Endpoints

28 AUGUST 2026·2 MIN READ·1 SOURCE·Trusted source

Google has launched Gemini 3.5 Transcribe, an API-only speech-to-text model featuring a 2.6% non-streaming word error rate and automatic code-switching support for over 85 languages.

Google Releases Gemini 3.5 Transcribe with 2.6% WER and Dual Endpoints

Key takeaways · 3

  • 01

    Achieves 2.6% average WER for pre-recorded audio and 4.0% for streaming tasks.

  • 02

    Improves time to final transcription by 70% compared to the previous Chirp 3 model.

  • 03

    Deployment is strictly via API; no open weights or self-hosted paths are available.

Capabilities and Architecture

Google has released Gemini 3.5 Transcribe, a new speech-to-text model designed for real-time voice interfaces and recorded audio. [1] The system provides automatic detection for more than 85 languages, which includes support for mid-sentence code-switching. [1] Performance metrics measured by Artificial Analysis show an average word error rate of 2.6% for non-streaming tasks and 4.0% for streaming tasks. [1] Furthermore, the model improves time to final transcription by 70% over the previous Chirp 3 model. [1]

Deployment and Endpoints

The model ships as two distinct endpoints: `gemini-3.5-transcribe` for pre-recorded files and `gemini-3.5-transcribe-live` for bidirectional streaming. [1] These two endpoints do not share the same pricing, limits, or feature sets. [1] The deployment is strictly API-only as a managed service, with no open weights available for self-hosting. [1] Currently, the developer and enterprise tracks remain in public preview. [1]

What it means

Google is targeting the enterprise transcription market by splitting Gemini 3.5 Transcribe into discrete batch and streaming endpoints. The 70% speed improvement over Chirp 3 positions it strongly for real-time customer experience applications. However, the lack of an open-weights release forces organizations into a managed-service ecosystem, contrasting with self-hosted open alternatives. What the sources don't address: How the specific pricing structures and rate limits compare between the live and batch endpoints.

Google's new speech-to-text release emphasizes speed and multi-language support for enterprise voice applications. The hard split between batch and streaming endpoints requires careful architecture planning for developers.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 28 August 2026

    Google Releases Gemini 3.5 Transcribe with 2.6% WER and Dual Endpoints

  2. 28 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.