Skip to main content

Google Launches Gemini 3.5 Transcribe and Omni 1.1 Flash for Audio and Video Workflows

27 AUGUST 2026·2 MIN READ·2 SOURCES·Official source plus independent coverage

Google released two new generative AI models to developers, offering advanced real-time speech transcription and enhanced generative video capabilities.

Google Launches Gemini 3.5 Transcribe and Omni 1.1 Flash for Audio and Video Workflows

Key takeaways · 2

  • 01

    Gemini 3.5 Transcribe handles background noise and complex jargon via Live and Interactions APIs.

  • 02

    Gemini Omni 1.1 Flash allows video scene extensions in 10-second increments up to 40 seconds.

Gemini 3.5 Transcribe

Gemini 3.5 Transcribe is a new speech-to-text model that converts raw audio inputs into highly accurate, formatted text while effectively handling background noise and complex industry jargon. [1] Developers can access the model via a Live API for real-time streaming applications, as well as an Interactions API for pre-recorded audio processing tasks. [1]

Gemini Omni 1.1 Flash

Google introduced Gemini Omni 1.1 Flash, which brings generative video capabilities to the Gemini API in Google AI Studio. [2] Omni 1.1 Flash allows users to extend videos in 10-second increments up to a cumulative length of 40 seconds. [2] The model also lets users specify starting and ending frames to generate continuous video between two keyframes. [2]

What it means

Google is expanding its developer ecosystem with production-ready multimodal tools, offering both sub-second audio transcription and high-resolution video generation. By releasing these capabilities concurrently via the Gemini API, Google enables developers to build complex applications ranging from voice agents to media editing software. This dual release positions Google's AI Studio as a comprehensive hub for both audio and visual generative tasks. What the sources don't address: the pricing structures and specific usage limits for processing large volumes of media through these new APIs.

These releases provide developers with sophisticated multimodal tools for enterprise applications. The models enable both low-latency voice agents and high-fidelity video production workflows.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 27 August 2026

    Google Launches Gemini 3.5 Transcribe and Omni 1.1 Flash for Audio and Video Workflows

  2. 27 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.