whisper(ml)x requires an Apple Silicon Mac (M1, M2, M3, or M4). MLX inference runs on the Apple Silicon GPU automatically — no configuration needed.
1
Install whisper(ml)x
Install the package using pip or uv:
2
Run your first transcription
Transcribe an audio file from the command line. The model downloads automatically from Hugging Face on first run.This produces transcript files in your current directory in all supported formats (SRT, VTT, TXT, TSV, JSON).
3
Transcribe with the Python API
Load a model and transcribe an audio file programmatically:
result is a TranscriptionResult with the following structure:4
Add alignment for word-level timestamps
Run forced alignment to get precise start and end times for every word:The aligned result is an
AlignedTranscriptionResult:Next steps
Installation
Set up speaker diarization and learn about model caching.
CLI reference
See all available flags for the
whispermlx command.Diarization guide
Assign speaker labels to segments and words.
Python API
Full reference for
load_model, align, and DiarizationPipeline.