Skip to main content

Whisper models

whisper(ml)x uses mlx-whisper for transcription. All inference runs on Apple Silicon’s GPU via the MLX framework.

Short name resolution

You can identify a model by its short name (e.g., "large-v3") or by a full HuggingFace repo ID. The _resolve_mlx_model() function in asr.py handles the mapping: if the string contains a /, it is used as-is; otherwise it is looked up in MLX_MODEL_MAP.
If you pass an unknown short name that is not in MLX_MODEL_MAP and does not contain a /, whisper(ml)x logs a warning and passes the string through unchanged. The model load will then fail if no matching repo exists on HuggingFace.

Model map

large and large-v1 resolve to the same repo. large-v3-turbo and turbo are also identical — turbo is a convenient alias.

Choosing a model

tiny / base

Fastest inference, lowest memory use. Best for quick experiments or real-time-adjacent use cases where accuracy is less critical.

small / medium

Good balance of speed and accuracy. Suitable for most production use cases where latency matters.

large-v3

Highest accuracy across languages. Use when quality is the priority and inference time is acceptable.

turbo (large-v3-turbo)

A distilled version of large-v3 that is significantly faster with only a minor accuracy tradeoff. A strong default choice for most use cases.
English-only variants (tiny.en, base.en, small.en, medium.en) are optimized for English and offer slightly better accuracy on English audio. Passing a non-English language to an .en model falls back to English automatically with a warning.

Model caching

Models are downloaded from HuggingFace on first use and cached in the standard HuggingFace cache directory (~/.cache/huggingface/hub by default). Subsequent runs load from cache and do not require network access. To force offline loading from cache, pass --model_cache_only on the CLI or local_files_only=True to load_model().

Alignment models

The alignment stage uses wav2vec2 models to produce word-level timestamps. Models are selected per language automatically from two sources:
  • torchaudio pipelines — used for en, fr, de, es, it
  • HuggingFace models — used for all other supported languages

Torch audio pipeline models

HuggingFace alignment models (selection)

30+ languages are supported in total. For the full list, see alignment.py. To use a custom alignment model, pass --align_model <HF_REPO_ID> on the CLI or the model_name argument to load_align_model().
If no default alignment model exists for your language, load_align_model() raises a ValueError. You must supply a custom wav2vec2 model via --align_model or skip alignment with --no_align.

Diarization model

The diarization pipeline uses pyannote. The default model is:
To use a different model, pass --diarize_model <MODEL_NAME> on the CLI or set model_name when constructing DiarizationPipeline directly. Diarization models are also cached in the HuggingFace cache directory. Some models require accepting terms of use on HuggingFace and providing an access token via --hf_token.