Meta Launches Muse Voice Transcribe for Real-Time, Multi-Speaker Audio
Meta has introduced Muse Voice Transcribe, a real-time audio perception model capable of transcribing and identifying more than 20 simultaneous speakers.
Key takeaways · 3
- 01
Muse Voice Transcribe identifies who is speaking among more than 20 individuals.
- 02
The tool consolidates transcription and turn-detection tasks that usually require separate systems.
- 03
The model detects mid-sentence language changes across its 25 validated launch languages.
Muse Voice Transcribe
Meta presented Muse Voice Transcribe, the first real-time audio perception model developed by Meta Superintelligence Labs. [1] The artificial intelligence transcribes words as a person speaks and can identify who said what among more than 20 speakers. [1] It also detects when an intervention ends, integrating functions that typically require separate systems into a single tool. [1]
The model is trained on more than 70 languages and recognizes language changes within the same sentence. [1] For its initial launch, Meta validated 25 languages, including Chinese, French, Hindi, and Spanish. [1] The artificial intelligence also has the capacity to pause in order to continue the transcription later. [1]
What it means
By integrating transcription, speaker identification, and turn-detection into one tool, Meta targets the difficulty of processing audio with more than 20 distinct speakers. This system enters a market alongside Google's Gemini 3.5, which recently launched with the ability to transcribe over 85 languages, but Muse Voice Transcribe attempts to distinguish itself by processing mid-sentence language switches and managing high-volume multi-speaker environments in real time. What the sources don't address: How the model's accuracy and latency degrade when scaling up to the 30-voice limit shown in its demonstration.
The ability to accurately transcribe and attribute dialogue in environments with more than 20 speakers simplifies complex audio processing. Integrating these capabilities into a single model reduces the need for chaining separate transcription and diarization systems.
Why it matters
Turn this story into practical AI skill after launch.
Get the release link for daily sessions built around your role and industry.
Join the waitlistHow this developed
3 September 2026
Meta Launches Muse Voice Transcribe for Real-Time, Multi-Speaker Audio
2 September 2026
Event created from source cluster.