Whisper models
whisper(ml)x uses mlx-whisper for transcription. All inference runs on Apple Silicon’s GPU via the MLX framework.Short name resolution
You can identify a model by its short name (e.g.,"large-v3") or by a full HuggingFace repo ID. The _resolve_mlx_model() function in asr.py handles the mapping: if the string contains a /, it is used as-is; otherwise it is looked up in MLX_MODEL_MAP.
If you pass an unknown short name that is not in
MLX_MODEL_MAP and does not contain a /, whisper(ml)x logs a warning and passes the string through unchanged. The model load will then fail if no matching repo exists on HuggingFace.Model map
large and large-v1 resolve to the same repo. large-v3-turbo and turbo are also identical — turbo is a convenient alias.Choosing a model
tiny / base
Fastest inference, lowest memory use. Best for quick experiments or real-time-adjacent use cases where accuracy is less critical.
small / medium
Good balance of speed and accuracy. Suitable for most production use cases where latency matters.
large-v3
Highest accuracy across languages. Use when quality is the priority and inference time is acceptable.
turbo (large-v3-turbo)
A distilled version of large-v3 that is significantly faster with only a minor accuracy tradeoff. A strong default choice for most use cases.
tiny.en, base.en, small.en, medium.en) are optimized for English and offer slightly better accuracy on English audio. Passing a non-English language to an .en model falls back to English automatically with a warning.
Model caching
Models are downloaded from HuggingFace on first use and cached in the standard HuggingFace cache directory (~/.cache/huggingface/hub by default). Subsequent runs load from cache and do not require network access.
To force offline loading from cache, pass --model_cache_only on the CLI or local_files_only=True to load_model().
Alignment models
The alignment stage uses wav2vec2 models to produce word-level timestamps. Models are selected per language automatically from two sources:- torchaudio pipelines — used for
en,fr,de,es,it - HuggingFace models — used for all other supported languages
Torch audio pipeline models
HuggingFace alignment models (selection)
30+ languages are supported in total. For the full list, see
alignment.py.
To use a custom alignment model, pass --align_model <HF_REPO_ID> on the CLI or the model_name argument to load_align_model().
Diarization model
The diarization pipeline uses pyannote. The default model is:--diarize_model <MODEL_NAME> on the CLI or set model_name when constructing DiarizationPipeline directly.
Diarization models are also cached in the HuggingFace cache directory. Some models require accepting terms of use on HuggingFace and providing an access token via --hf_token.