Requirements
- Apple Silicon Mac — M1, M2, M3, or M4 chip
- macOS — any version supported by your chip
- Python — 3.10, 3.11, 3.12, or 3.13
Install whisper(ml)x
mlx-whisper— ASR inference on Apple Silicontorchandtorchaudio— CPU wheels, used by VAD and alignment modelstransformers— wav2vec2 alignment modelspyannote-audio— VAD and speaker diarization
torch and torchaudio are installed as CPU wheels from the PyTorch CPU index. They are used for VAD and alignment — not for Whisper inference, which runs via MLX.Model downloads
Models are downloaded from Hugging Face on first use and cached at~/.cache/whisper by default. You do not need to download anything manually.
Short model names are automatically resolved to their mlx-community equivalents:
You can also pass a full Hugging Face repo ID directly:
--model_dir on the CLI or set download_root in load_model().
Speaker diarization setup
Speaker diarization requires a Hugging Face access token and acceptance of the pyannote model agreement.1
Create a Hugging Face access token
Go to https://huggingface.co/settings/tokens and create a token with read access.
2
Accept the model agreement
Visit the pyannote/speaker-diarization-community-1 model page and accept the user agreement while logged in to your Hugging Face account.
3
Pass the token to whisper(ml)x
Provide your token via the CLI flag or the Python API: