Skip to main content
Voice Activity Detection (VAD) is the first stage of the whisper(ml)x pipeline. Before any transcription happens, the VAD model scans the full audio and identifies the regions that contain speech.

What VAD does

Running Whisper on silence or background noise causes it to hallucinate text. VAD solves this by:
  1. Scanning the full audio waveform and producing a list of {start, end} speech segments.
  2. Merging adjacent segments into chunks no longer than chunk_size seconds.
  3. Passing only those chunks to Whisper — skipping silence entirely.
This reduces transcription errors on quiet or noisy recordings and lowers inference time by not processing non-speech regions.

Two VAD backends

whisper(ml)x supports two VAD backends. Both implement the same abstract Vad interface and produce the same segment format.

Pyannote (default)

Loads a local model from whispermlx/assets/pytorch_model.bin. More accurate on a wide range of audio types, including noisy environments. Uses hysteresis thresholding with a min-cut operation to avoid creating segments longer than chunk_size.

Silero

Loads via torch.hub from snakers4/silero-vad. Lighter weight with no bundled assets. Requires network access on first use. Works well for clean recordings.
The device parameter in load_model() controls where the VAD model runs, not MLX Whisper. For Apple Silicon, "cpu" is recommended for VAD. MLX inference always runs on the GPU automatically.

Key parameters

vad_onset and vad_offset implement hysteresis: a higher onset prevents false starts, while a lower offset allows the model to continue a speech segment through brief dips in confidence.
For difficult audio — noisy environments, quiet speakers, or recordings with long pauses — try reducing vad_onset to 0.3–0.4 to catch more speech, and vad_offset to 0.2–0.3 to avoid cutting off words at the end of sentences.

Python API

Pass VAD configuration to load_model() via the vad_method and vad_options arguments:
You can also pass a pre-instantiated VAD model via the vad_model argument. When vad_model is set, vad_method is ignored:

CLI

Use the --vad_method, --vad_onset, --vad_offset, and --chunk_size flags:
To use pyannote (the default), you can omit --vad_method entirely: