Is Meta's Muse Voice Transcribe Really SOTA?

Verified

Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, 2026, with CEO Mark Zuckerberg calling it 'SOTA in streaming speech-to-text.' Independent benchmark firm Artificial Analysis confirms the model ranks #1 on its AA-WER Streaming test at 3.1% word error rate, ahead of Cartesia Ink-2 (3.4%), ElevenLabs Scribe v2 (3.6%), OpenAI's GPT Live Transcribe (3.9%), and Google's Gemini 3.5 Transcribe Live (4%). The lead is real but narrow, and the test measures English only.

The claim

"Meta's Muse Voice Transcribe model is state-of-the-art (SOTA) in streaming speech-to-text, beating models from OpenAI, Google, Cartesia, and ElevenLabs." — Meta AI Research, 2026-09-01

Frequently asked questions

What benchmark supports Meta's 'SOTA' claim?

Artificial Analysis's AA-WER Streaming test, which scores word error rate for real-time transcription. Muse Voice Transcribe recorded 3.1% WER, the lowest of any model the firm tracks.

Does the SOTA claim cover all 70-plus languages Meta trained on?

No. The benchmark backing the claim is English-only. Meta says it trained on more than 70 languages but had extensively validated only 25 of them at launch.

Is Muse Voice Transcribe also state-of-the-art at identifying individual speakers?

Not by the same margin. Meta leads public speaker-diarization benchmarks, but at a 17.5% error rate — a level industry coverage describes as still far from solved across the field.

Reviewed by the AI Ledger editorial team · Last updated 2026-09-03. Verdicts follow our published methodology; spot an error? tell us and we'll re-review it.