TECH Signal 309
Meta launches Muse Voice Transcribe handling 20+ speakers and mid-sentence code-switching across 70+ languages
Meta Superintelligence Lab released Muse Voice Transcribe, a real-time speech-to-text model that performs speaker diarization, endpointing, and multilingual transcription in a single model.
The model is available now via Meta's Model API at $3 per 1,000 audio minutes and already powers dictation in the Meta AI Mac app and Muse Code, giving developers a concrete alternative to Google's Gemini 3.5 Transcribe released days earlier. The single-model approach to diarization and endpointing could simplify pipelines that currently stitch together separate components for speaker identification and transcription.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Muse Voice Transcribe handles 20+ speakers, 70+ training languages with 25 validated at launch, and hour-long sessions natively in one model.
The model uses adaptive delay, waiting longer on difficult words and committing faster on easy ones to balance accuracy and latency.
It is accessible via Meta's Model API at $3 per 1,000 audio minutes, with a demo on Meta's research blog, but Meta has not announced integration into its flagship consumer services.
THE CLUSTER
↗