Skip to main content
Loads a Whisper model and returns an MLXWhisperPipeline ready for transcription. The pipeline bundles a VAD model that segments audio before each chunk is passed to MLX Whisper on Apple Silicon.

Parameters

string
required
Short model name (e.g. "large-v3", "small") or a full Hugging Face repo ID (e.g. "mlx-community/whisper-large-v3-mlx"). Short names are resolved via a built-in map:English-only variants (tiny.en, base.en, small.en, medium.en) are also available. Pass a full Hugging Face repo ID for any other MLX-compatible model.
string
required
Device for the VAD model. Use "cpu" or "mps". MLX Whisper inference always runs on Apple Silicon GPU automatically — this parameter does not affect it.
integer
default:"0"
Kept for API compatibility. Unused by MLX inference.
string
default:"default"
Kept for API compatibility. Ignored by MLX inference.
object
Dict of ASR options. Only the "initial_prompt" key is used — all other keys are ignored.
string
BCP-47 language code (e.g. "en", "fr", "ja"), or None to enable automatic language detection from the first audio chunk.
Vad
A pre-instantiated VAD model. When provided, vad_method is ignored. Use this to share a single VAD instance across multiple calls.
string
default:"pyannote"
VAD backend to use when vad_model is not provided. Accepted values: "pyannote" (default) or "silero".
object
Overrides for VAD configuration. Keys and defaults:
string
default:"transcribe"
Default task for the pipeline. "transcribe" produces output in the source language; "translate" produces an English translation. Can be overridden per-call on .transcribe().
boolean
default:"false"
When True, use only locally cached model files and do not attempt to download from Hugging Face. Raises an error if the model is not cached.
integer
default:"4"
Kept for API compatibility. Ignored by MLX inference.

Returns

An MLXWhisperPipeline instance. Call .transcribe() on it to transcribe audio.

Full signature

Usage

The device parameter only controls where the VAD model (pyannote or silero) runs. MLX Whisper inference executes on the Apple Silicon GPU automatically regardless of this setting.

MLXWhisperPipeline.transcribe

Transcribe audio with the loaded model.

Voice Activity Detection

Learn about the VAD stage and available backends.