Why use whisper(ml)x
MLX inference
Runs transcription on the Apple Silicon GPU via unified memory — no CUDA, no discrete GPU required.
Word-level timestamps
Forced alignment via wav2vec2 produces precise start and end times for every word in the transcript.
Speaker diarization
Identify who spoke when using pyannote-audio’s speaker diarization pipeline.
VAD preprocessing
Voice activity detection (pyannote or silero) strips silence before transcription, reducing inference time on long recordings.
Architecture overview
whisper(ml)x runs a three-stage pipeline on your audio:1
VAD — Voice activity detection
Audio is preprocessed by pyannote or silero VAD to detect speech segments. Silence is discarded and speech chunks are merged into 30-second windows before being passed to the ASR stage.
2
ASR — Transcription
Each speech segment is transcribed by mlx-whisper, which runs natively on the Apple Silicon GPU. The result is a list of
SingleSegment dicts containing start, end, and text.3
Alignment and diarization (optional)
A wav2vec2 alignment model attaches word-level timestamps to each segment, producing
SingleAlignedSegment dicts. Optionally, a pyannote diarization model assigns speaker labels to each word.System requirements
- Apple Silicon Mac — M1, M2, M3, or M4 chip (required for MLX inference)
- macOS — any version supported by your chip
- Python — 3.10, 3.11, 3.12, or 3.13
MLX only runs on Apple Silicon. whisper(ml)x will not work on Intel Macs or non-macOS platforms for transcription. The
device parameter controls PyTorch models (VAD, alignment, diarization) only — MLX Whisper always uses the Apple Silicon GPU automatically.Get started
Quick start
Transcribe your first audio file in under five minutes.
Installation
Install with pip or uv and configure speaker diarization.
CLI reference
Explore every command-line flag and option.
Python API
Integrate transcription, alignment, and diarization into your code.