Meta Launches Muse Voice Transcribe for Real-Time Multilingual Transcription
The model supports 20 speakers and 70 languages, with 25 validated at launch. It is available through the Meta Model API at $3 per 1,000 audio-minutes.
Meta Superintelligence Labs has launched Muse Voice Transcribe, a real-time audio perception model capable of streaming automatic speech recognition, speaker diarization, and endpointing. The model is designed to transcribe speech as it happens, separate multiple speakers across recordings, and identify when a speaker has finished talking, all without requiring post-processing. This marks Meta's first foray into real-time transcription and is the latest release from the Meta Superintelligence Lab (MSL).
Muse Voice Transcribe is integrated into Meta AI for Mac, Muse Code, and available to developers through the Meta Model API. The model was trained across more than 70 languages, with 25 of those languages validated at launch. It is capable of handling dictation and transcription for more than 20 speakers simultaneously, making it suitable for a wide range of applications, from virtual meetings to transcription services.
Meta claims that Muse Voice Transcribe ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026. The model is available today through the Meta Model API for $3 per 1,000 audio-minutes, which equates to $0.18 per hour. This pricing structure is intended to make the model accessible to developers and businesses looking to integrate real-time transcription into their applications.
The introduction of Muse Voice Transcribe may influence the broader AI transcription market, potentially increasing competition and prompting other companies to enhance their own transcription models. The model's capabilities could lead to increased adoption of real-time transcription in various industries, including healthcare, education, and customer service. However, the cost and potential vendor lock-in associated with using the Meta Model API may be factors for organizations considering integration.
As the model continues to develop, its impact on the AI transcription landscape will depend on its performance, scalability, and the extent to which it can be adapted to different use cases. The release of Muse Voice Transcribe represents a significant step forward for Meta in the field of real-time audio processing and may set new standards for transcription accuracy and speed.
Sources
- https://9to5mac.com/2026/09/01/meta-launches-muse-voice-transcribe-for-real-time-voice-dictation-on-mac/
- https://indianexpress.com/article/technology/artificial-intelligence/meta-launches-muse-voice-transcribe-with-support-for-5-indian-languages-10859850/
- https://www.engadget.com/2249112/meta-new-ai-transcription-model-can-distinguist-between-multiple-speakers-and-languages-in-real-time/
- https://www.gadgets360.com/ai/news/meta-muse-voice-transcribe-supports-5-indian-languages-more-than-70-global-11991706
- https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing/