Skip to main content
whisper(ml)x requires an Apple Silicon Mac (M1, M2, M3, or M4). MLX inference runs on the Apple Silicon GPU automatically — no configuration needed.
1

Install whisper(ml)x

Install the package using pip or uv:
Use uv for faster installs and automatic virtual environment management.
2

Run your first transcription

Transcribe an audio file from the command line. The model downloads automatically from Hugging Face on first run.
This produces transcript files in your current directory in all supported formats (SRT, VTT, TXT, TSV, JSON).
Use --model large-v3 for the best accuracy. It maps to mlx-community/whisper-large-v3-mlx automatically.
3

Transcribe with the Python API

Load a model and transcribe an audio file programmatically:
result is a TranscriptionResult with the following structure:
4

Add alignment for word-level timestamps

Run forced alignment to get precise start and end times for every word:
The aligned result is an AlignedTranscriptionResult:

Next steps

Installation

Set up speaker diarization and learn about model caching.

CLI reference

See all available flags for the whispermlx command.

Diarization guide

Assign speaker labels to segments and words.

Python API

Full reference for load_model, align, and DiarizationPipeline.