Skip to content
Muse Atlas
Explore Muse

Muse Voice

From spoken words to useful text.

Concept illustration of Muse Voice

Speech in. Text out.

Muse Voice Transcribe is Meta's speech-to-text model. Applications access it through Model API, an interface for sending audio and receiving text, with streaming and file transcription, speaker labels and detection of speech turns.

What Voice is good at

Transcribe as people speak

Use streaming speech input in an application.

Work with recordings

Turn audio files into transcripts for review.

Follow speakers and turns

Use speaker labels and turn-level timing where supported.

Turn a conversation into an actionable draft.

See where transcription ends and application reasoning begins.

Illustrative scenario
Speaker 1
Can we review the prototype tomorrow?

Speech is converted into text.

Speaker 2
Yes. Let’s make it ten in the morning.

A new turn can be attributed to another speaker.

Application
Proposed task: review prototype tomorrow at 10:00.

An application can interpret the transcript. Voice Transcribe itself does not schedule the meeting.

Illustrative transcript. No microphone recording or live transcription.

Where it fits in Muse

Voice Transcribe turns audio into text. Spark or your application can interpret that text and choose an action. Voice Transcribe itself does not generate speech.

Real examples

Technical architecture

Audio is sent to Meta Model API over realtime WebSocket or file upload endpoints. The model returns transcript text with turn-level timing and optional speaker labels.

  1. 1Audio input
  2. 2Voice Transcribe
  3. 3Text + turns
  4. 4Your application

Specifications and limitations

Model ID
muse-voice-transcribe-1.0
API pricing — checked September 29, 2026
$0.18 per hour of processed audio; streaming and file transcription share the rate. Check current terms in the footer resources.
Availability
Meta Model API

Keep in mind

  • Speech-to-text only — no text-to-speech in this model
  • Turn-level timestamps, not word-level
  • Separate from Muse realtime conversational voice and avatar experiences