What VAD does
Running Whisper on silence or background noise causes it to hallucinate text. VAD solves this by:- Scanning the full audio waveform and producing a list of
{start, end}speech segments. - Merging adjacent segments into chunks no longer than
chunk_sizeseconds. - Passing only those chunks to Whisper — skipping silence entirely.
Two VAD backends
whisper(ml)x supports two VAD backends. Both implement the same abstractVad interface and produce the same segment format.
Pyannote (default)
Loads a local model from
whispermlx/assets/pytorch_model.bin. More accurate on a wide range of audio types, including noisy environments. Uses hysteresis thresholding with a min-cut operation to avoid creating segments longer than chunk_size.Silero
Loads via
torch.hub from snakers4/silero-vad. Lighter weight with no bundled assets. Requires network access on first use. Works well for clean recordings.The
device parameter in load_model() controls where the VAD model runs, not MLX Whisper. For Apple Silicon, "cpu" is recommended for VAD. MLX inference always runs on the GPU automatically.Key parameters
vad_onset and vad_offset implement hysteresis: a higher onset prevents false starts, while a lower offset allows the model to continue a speech segment through brief dips in confidence.
Python API
Pass VAD configuration toload_model() via the vad_method and vad_options arguments:
vad_model argument. When vad_model is set, vad_method is ignored:
CLI
Use the--vad_method, --vad_onset, --vad_offset, and --chunk_size flags:
--vad_method entirely: