Skip to main content
whisper(ml)x transcribes audio by running voice activity detection (VAD) to segment speech, then passing each chunk through an MLX Whisper model running on the Apple Silicon GPU.

CLI usage

The simplest way to transcribe is to pass an audio file and a model name:
Specify the language, task, and output directory explicitly:
Transcribe multiple files at once by listing them as positional arguments:

Python API

Load a model with whispermlx.load_model, then call .transcribe() on it:
The device parameter controls where the VAD model runs. MLX Whisper inference always uses the Apple Silicon GPU automatically — you do not need to configure it.

Available models

Pass a short model name or a full Hugging Face repo ID: English-only variants (tiny.en, base.en, small.en, medium.en) are also available.

Transcription options

Language detection

By default whisper(ml)x detects the spoken language automatically from the first audio chunk. Set --language (CLI) or pass language= to load_model or .transcribe() to skip detection:

Task

The task parameter controls whether to produce a transcript in the source language or an English translation:
When task is set to translate, alignment is skipped automatically because word-level alignment cannot be performed on translated text.

Batch size

The --batch_size flag is accepted for API compatibility but has no effect on MLX inference — the pipeline processes VAD segments sequentially regardless of this value.

Chunk size

chunk_size sets the maximum length (in seconds) of each VAD segment before it is split. The default is 30 seconds:

Initial prompt

Provide a text prompt to bias the first transcription window. Useful for guiding punctuation style or domain vocabulary:

Hotwords

Hint phrases for rare or technical terms that the model might otherwise mishear:

Progress tracking

Pass a progress_callback to .transcribe() to receive progress updates as a percentage:
The callback receives a float between 0 and 100 after each VAD segment is processed.

TranscriptionResult structure

model.transcribe() returns a TranscriptionResult TypedDict:
Each SingleSegment contains:
After running alignment, segments are upgraded to SingleAlignedSegment which adds a words list. See the alignment guide for details.

Next steps

Word-level alignment

Add precise word timestamps to your transcript.

Speaker diarization

Label each segment with the speaker who said it.

Output formats

Export to SRT, VTT, JSON, and more.

CLI reference

Browse every available CLI flag.