Skip to main content

Microsoft Launches Three MAI Models for Real-Time Transcription and Voice

2 OCTOBER 2026·2 MIN READ·1 SOURCE·Trusted source

Microsoft has expanded its MAI audio lineup with a low-latency streaming transcription model and two newly named voice models.

Microsoft Launches Three MAI Models for Real-Time Transcription and Voice

Key takeaways · 3

  • 01

    Evaluate MAI-Transcribe-2-Streaming when real-time transcripts and low-latency operation are core requirements.

  • 02

    Treat MAI-Voice-2.1 and its Flash variant as distinct options pending further technical details.

  • 03

    Request availability, pricing, language support, and performance data before selecting any model for production.

Three New Audio Models

Microsoft launched MAI-Transcribe-2-Streaming, a new model for low-latency, real-time transcripts. [1] The release describes MAI-Transcribe-2-Streaming specifically as a streaming transcription model. [1] Microsoft also launched two new voice models in the same announcement. [1]

Those voice models are named MAI-Voice-2.1 and MAI-Voice-2.1-Flash. [1] The announced lineup therefore contains one model focused on real-time transcripts and two distinct, separately named voice models. [1]

What it means

Microsoft is presenting streaming transcription and two voice models in the same announcement, placing real-time transcripts and voice capabilities side by side. MAI-Transcribe-2-Streaming has an explicit low-latency, real-time positioning, while the source supplies only the names of MAI-Voice-2.1 and MAI-Voice-2.1-Flash. That imbalance makes the transcription model’s intended role clearer than the differences between the voice models. What the sources don't address: how the two voice models differ in performance, pricing, availability, or supported languages.

The release gives AI practitioners a Microsoft model explicitly positioned for low-latency, real-time transcription, alongside two new voice options. Production evaluations will still require details about model access, performance, pricing, and language coverage.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 2 October 2026

    Microsoft Launches Three MAI Models for Real-Time Transcription and Voice

  2. 2 October 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.