Skip to main content

Microsoft launches streaming transcription model in public preview

3 OCTOBER 2026·2 MIN READ·4 SOURCES

Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026, describing it as its first streaming transcription model. The model converts live speech to text and is available in public preview, with Microsoft listing its price at $0.54 per audio hour through December 31, 2026.

Microsoft launches streaming transcription model in public preview

Key takeaways · 4

  • 01

    Developers can access the model on Microsoft Foundry and integrate it using an OpenAI Realtime-compatible API over WebSockets.

  • 02

    Microsoft lists the price at $0.54 per audio hour through December 31, 2026.

  • 03

    The model supports 60 languages and is designed for varied accents, speaking styles and real-world audio conditions.

  • 04

    Treat the current release as an evaluation opportunity: it is in public preview, without an SLA, and Microsoft Learn does not recommend it for production workloads.

A model built for live speech

Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026, and describes it as its first streaming transcription model, built in-house by the Microsoft AI team.[1][2][3] The model is designed to turn live audio into text as it is spoken.[3] It continuously processes audio, returning fast partial hypotheses before stable final results.[3] Microsoft says it supports 60 languages and is designed for varied accents, speaking styles and real-world audio conditions.[3] Listed uses include live captions, voice agents, dictation, call experiences and other interactive products.[3]

Accuracy and price

Microsoft lists the model's final word error rate at 2.5% and says it recorded a lower rate than the competitors in its Artificial Analysis comparison.[3] That comparison lists Grok Transcribe 2.0 at 2.7% average word error rate, compared with 2.5% for Microsoft's model.[3] SiliconReport reports that Microsoft cited a 2.5% final word-error rate across eight hours of test audio for the streaming benchmark.[1] Microsoft lists a price of $0.54 per audio hour through December 31, 2026.[3] The comparison is limited to the competitors Microsoft lists; the evidence does not establish a broader ranking.

Developer access and integration

SiliconReport says the model is available on Microsoft Foundry.[1] Microsoft's model card says developers can integrate it through an OpenAI Realtime-compatible API over WebSockets.[3] Microsoft Learn says both the Realtime API and Azure Speech SDK integration methods support the model and return intermediate and final transcription results.[4] SiliconReport reports that Microsoft also made two text-to-speech models available alongside the transcription model, and says the three models are available on Microsoft Foundry, Vercel, Azure Voice Live and MAI Playground, with LiveKit integration planned.[1] Microsoft lists East US 2 as a region where availability is coming soon.[3]

Preview caveats and intended uses

Microsoft Learn identifies the model as being in public preview and says the preview has no service-level agreement and is not recommended for production workloads.[4] Those terms matter for teams deciding whether to test the model or depend on it in a live service. SiliconReport reports that Microsoft is targeting contact center agents, multilingual assistants and interactive learning tools.[1] It also reports that the model produces initial partial hypotheses in just over 100 milliseconds and stable transcripts in 0.13 seconds.[1] Microsoft says continuous language detection lets voice agents prepare tool calls and reason mid-utterance.[1]

For teams building speech-enabled products, the release offers a streaming option with stated support for multiple languages and developer integration paths. The public-preview status, absence of an SLA and production-workload warning make testing and deployment readiness separate decisions.

Why it matters
Story quiz

Test yourself on this story — 1 question.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 3 October 2026

    Microsoft launches streaming transcription model in public preview

Sources

AI fluency, one session a day, built for your work.