Skip to main content

Meta Releases Muse Voice Transcribe API for Real-Time Audio Processing

2 SEPTEMBER 2026·2 MIN READ·2 SOURCES·Trusted source

Meta Superintelligence Labs has launched Muse Voice Transcribe, a single autoregressive model that combines streaming ASR, speaker diarization, and endpointing.

Meta Releases Muse Voice Transcribe API for Real-Time Audio Processing

Key takeaways · 3

  • 01

    Muse Voice Transcribe replaces three-part audio systems with a single autoregressive model.

  • 02

    The model handles diarization for over 20 speakers simultaneously.

  • 03

    It is only available via an API, with no open weights released for self-hosting.

Model Architecture

Meta Superintelligence Labs announced Muse Voice Transcribe, which combines streaming ASR, speaker diarization for over 20 speakers, and endpointing into a single autoregressive model. [1][2]

The system replaces traditional voice stacks that use three separate systems for transcribing, separating speakers, and detecting when a user stops talking. [1] As part of the Muse Spark family, the multimodal model receives audio in 80ms chunks at 12.5 Hz. [1]

Availability and Cost

The model is available exclusively as a hosted API via the Meta Model API under the name muse-voice-transcribe-1.0. [1] Meta has not released the weights, meaning there is no path for self-hosting. [1]

The API costs $3.00 per 1,000 audio minutes, which equals $0.18 per hour. [1] The system already powers dictation for Muse Code and Meta AI for Mac. [1]

What it means

By collapsing three separate audio processing jobs into one autoregressive model, Meta removes the traditional hand-offs that typically introduce latency and new failure modes into voice stacks. Because audio chunks are transformed directly into soft tokens without separate alignment stages, real-time transcription is streamlined. However, the decision to withhold the model weights limits its reach, requiring enterprise users to rely entirely on Meta's infrastructure rather than deploying the system in secure, self-hosted environments. What the sources don't address: How the model's raw transcription accuracy compares to existing state-of-the-art open-source ASR models on standard benchmarks.

The transition from multi-component voice processing stacks to single autoregressive models reduces latency for real-time applications. However, the lack of open weights restricts deployment options for privacy-conscious organizations.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 2 September 2026

    Meta Releases Muse Voice Transcribe API for Real-Time Audio Processing

  2. 2 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.