Is Meta's Muse Voice Transcribe Really SOTA?
VerifiedMeta Superintelligence Labs launched Muse Voice Transcribe on September 1, 2026, with CEO Mark Zuckerberg calling it 'SOTA in streaming speech-to-text.' Independent benchmark firm Artificial Analysis confirms the model ranks #1 on its AA-WER Streaming test at 3.1% word error rate, ahead of Cartesia Ink-2 (3.4%), ElevenLabs Scribe v2 (3.6%), OpenAI's GPT Live Transcribe (3.9%), and Google's Gemini 3.5 Transcribe Live (4%). The lead is real but narrow, and the test measures English only.
Frequently asked questions
What benchmark supports Meta's 'SOTA' claim?
Artificial Analysis's AA-WER Streaming test, which scores word error rate for real-time transcription. Muse Voice Transcribe recorded 3.1% WER, the lowest of any model the firm tracks.
Does the SOTA claim cover all 70-plus languages Meta trained on?
No. The benchmark backing the claim is English-only. Meta says it trained on more than 70 languages but had extensively validated only 25 of them at launch.
Is Muse Voice Transcribe also state-of-the-art at identifying individual speakers?
Not by the same margin. Meta leads public speaker-diarization benchmarks, but at a 17.5% error rate — a level industry coverage describes as still far from solved across the field.