Microsoft Launches Three MAI Models for Real-Time Transcription and Voice
Microsoft has expanded its MAI audio lineup with a low-latency streaming transcription model and two newly named voice models.

Key takeaways · 3
- 01
Evaluate MAI-Transcribe-2-Streaming when real-time transcripts and low-latency operation are core requirements.
- 02
Treat MAI-Voice-2.1 and its Flash variant as distinct options pending further technical details.
- 03
Request availability, pricing, language support, and performance data before selecting any model for production.
Three New Audio Models
Microsoft launched MAI-Transcribe-2-Streaming, a new model for low-latency, real-time transcripts. [1] The release describes MAI-Transcribe-2-Streaming specifically as a streaming transcription model. [1] Microsoft also launched two new voice models in the same announcement. [1]
Those voice models are named MAI-Voice-2.1 and MAI-Voice-2.1-Flash. [1] The announced lineup therefore contains one model focused on real-time transcripts and two distinct, separately named voice models. [1]
What it means
Microsoft is presenting streaming transcription and two voice models in the same announcement, placing real-time transcripts and voice capabilities side by side. MAI-Transcribe-2-Streaming has an explicit low-latency, real-time positioning, while the source supplies only the names of MAI-Voice-2.1 and MAI-Voice-2.1-Flash. That imbalance makes the transcription model’s intended role clearer than the differences between the voice models. What the sources don't address: how the two voice models differ in performance, pricing, availability, or supported languages.
The release gives AI practitioners a Microsoft model explicitly positioned for low-latency, real-time transcription, alongside two new voice options. Production evaluations will still require details about model access, performance, pricing, and language coverage.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
2 October 2026
Microsoft Launches Three MAI Models for Real-Time Transcription and Voice
2 October 2026
Event created from source cluster.