In his review of Google’s August 26 launch, Ryan Whitwam tested Gboard’s Rambler and watched Gemini 3.5 Transcribe remove verbal stumbles and turn spoken corrections into polished text. The senior technology reporter found the cleanup useful for short blocks of dictation.

The transcript was no longer exactly what he had said.

Muse Voice Transcribe, launched September 1 by Meta Superintelligence Labs, takes a different approach to real-time speech: one model transcribes words, separates speakers, and detects when a person has stopped talking.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

A narrow lead in English

Meta’s strongest evidence comes from the AA-WER Streaming test operated by Artificial Analysis. Word error rate is the percentage of words a transcript gets wrong compared with a verified reference transcript.

On the benchmark displayed September 1, Muse produced a 3.1% word error rate on the final transcript 0.16 seconds after speech ended. Cartesia’s Ink-2, in its semantic-endpoints configuration, followed at 3.4%, ElevenLabs’ Scribe v2 Realtime at 3.6%, OpenAI’s GPT Live Transcribe at 3.9%, and Google’s Gemini 3.5 Transcribe Live at 4%.

The result supports Meta’s claim that Muse ranks first, within the limits of this test. The lead over Cartesia is 0.3 percentage points. Artificial Analysis builds the index from about eight hours of English audio drawn from AA-AgentTalk, VoxPopuli, and Earnings22. It does not test Meta’s performance across the more than 70 languages used in training, including the 25 languages Meta had extensively verified at launch.

The earlier, less-settled transcript produces a tighter speed comparison. On the first partial transcript after speech ended, Muse recorded 3.6% error at 0.13 seconds in the September 1 results, slightly ahead of ElevenLabs. Cartesia Ink-2 with external endpoint detection reached 4% at 0.07 seconds, while Cartesia’s semantic endpoint setting returned 4.9% at 0.17 seconds. That cut gives Cartesia the faster response and Muse the more accurate text.

How Muse decides to wait

Muse processes incoming speech in 80-millisecond chunks. At each chunk, it can emit a word or keep listening. Meta calls this adaptive delay: the model waits for more context when a word is difficult and commits sooner when the audio is clear.

Reinforcement learning sets that behavior by combining a reward for lower word error with a reward for shorter delay. The two rewards are multiplied, making a choice poor if it performs badly on either measure. When the speaker stops, an empty-audio signal tells the model to release any text it is still holding.

The same model also emits markers for a speaker change and for the end of speech. At its September 1 launch, Muse supported audio longer than an hour and diarization for more than 20 speakers without a separate processing stage. It became available through the Meta Model API and already powered system-wide dictation in Meta AI for Mac, where holding the Fn key sends speech into any application, as well as Muse Code.

Speaker labels remain error-prone

Diarization is Muse’s weak column. Meta said the model led public diarization benchmarks at launch with a 17.5% error rate.

The language claim carries a separate limit. Meta trained Muse on more than 70 languages and recommended 25 verified languages on September 1, but the leaderboard behind its accuracy claim measures English only. A team choosing Muse for a multilingual call center cannot use the 3.1% figure as evidence that Spanish, Hindi, or German will perform the same way.

Google also shows why accuracy can depend on the job being measured. Gemini 3.5 Transcribe launched six days before Muse with a 5.50% live-speech word error rate on Google’s multilingual FLEURS evaluation, down from 7.32% for Chirp 3. Its streaming score on Artificial Analysis was 4% on September 1. But Google’s model is designed to remove filler words and resolve self-corrections such as changing a meeting from Tuesday to Wednesday.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

That editing can be helpful when the goal is a clean email and risky when the exact wording matters. Whitwam wrote that “you’re relying on the AI to accurately get the gist of your speech.” His test found that the model technically changed what he said, a poor fit for records where a disfluency or correction belongs in the transcript.

Meta’s price advantage

Muse costs $3 per 1,000 audio minutes through the Meta Model API as of its September 1 launch, equal to $0.18 an hour. Cartesia Ink-2 cost $4 per 1,000 minutes on the same date. ElevenLabs Scribe v2 Realtime and Deepgram Flux each cost $6.50 per 1,000 minutes, more than twice Meta’s rate.

The gap is larger against Google Cloud Speech-to-Text’s standard tier, priced around $0.96 an hour at launch. Meta’s $0.18 rate is $0.78 lower. For an application processing 1,000 hours, the listed rates imply $180 with Muse and about $960 with Google Cloud before any other charges.

Different tests produce different leaders

AssemblyAI’s own benchmark suite offers the clearest warning against treating one leaderboard as a final ranking. Its pre-recorded test, last updated July 17, covers synthetic medical audio, accented English from India, general speech, and webinars. AssemblyAI Universal-3.5 Pro led that table with a 4.35% average normalized word error rate, ahead of ElevenLabs Scribe V2 at 5.87% and Deepgram Nova-3 at 6.66%.

Muse does not appear in that July table, and AssemblyAI has a commercial interest in the result. The suite still exposes the central problem with any claim to be the best: its data differs from Artificial Analysis’s English streaming mix, and its pre-recorded models receive more context than a live system. A medical scribe, a meeting assistant, and a voice agent can fail on different words even when their average scores look close.

Meta has the lowest error rate on the September 1 AA-WER Streaming table and charges less than Cartesia, ElevenLabs, Deepgram, and Google at the listed rates. Will that combination hold when Muse is tested on multilingual calls, specialized names, and a competitor’s datasets?

Frequently Asked Questions

What is Muse Voice Transcribe?

It's Meta Superintelligence Labs' first real-time audio perception model, launched September 1, 2026, combining streaming speech recognition, speaker diarization for 20+ speakers, and endpointing (detecting when someone stops talking) in one model.

How accurate is Muse Voice Transcribe compared to rivals?

On Artificial Analysis's independent AA-WER Streaming benchmark, Muse scored a 3.1% word error rate, ahead of Cartesia Ink-2 (3.4%), ElevenLabs Scribe v2 Realtime (3.6%), OpenAI's GPT Live Transcribe (3.9%), and Google's Gemini 3.5 Transcribe Live (4%). The lead over the next-best model is 0.3 percentage points, and the test measures English only.

How much does Muse Voice Transcribe cost?

$3 per 1,000 audio minutes, or $0.18 per hour, through the Meta Model API. That's about 80% below Google Cloud Speech-to-Text's standard $0.96-per-hour rate, and cheaper than Cartesia Ink-2 ($4 per 1,000 minutes) and ElevenLabs Scribe v2 Realtime or Deepgram Flux ($6.50 per 1,000 minutes each).

What is Muse Voice Transcribe's biggest weakness?

Speaker diarization. Meta's own reported error rate for identifying who is speaking is 17.5%, even on the benchmark where it claims the lead, meaning speaker labeling is still far from solved.

Is Muse Voice Transcribe definitively the best speech-to-text model?

Not by every measure. A separate benchmark suite published by AssemblyAI, built on different test data, ranks its own Universal-3.5 Pro model first, showing that "best" depends heavily on which leaderboard and dataset is used.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

A Serious AI Workbench Starts With Ten Local MCP Servers
Deepgram Launches Flux Multilingual Speech Model With 10-Language Mid-Call Switching
Deepgram today announced Flux Multilingual, a conversational speech recognition model that supports 10 languages with real-time language detection and the ability to switch languages during an active
Google Expands Meet Speech Translation From 5 Languages to 70 With Gemini 3.5
Google released Gemini 3.5 Live Translate on Tuesday, an audio model that translates speech continuously across more than 70 languages instead of waiting for a speaker to finish. The model began rolli
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai