Skip to main content
whisper(ml)x processes audio through up to four sequential stages. Each stage builds on the output of the previous one, and later stages are optional depending on what you need.
1

Stage 1 — Voice Activity Detection (VAD)

The pipeline begins by identifying which portions of the audio actually contain speech. Raw audio is loaded at 16 kHz and passed to the VAD model, which returns a list of speech segments with start and end timestamps.Silence, noise, and non-speech regions are excluded before any transcription happens. This prevents Whisper from hallucinating text on silent sections and reduces the total amount of audio that needs to be processed.Adjacent VAD segments are merged into chunks up to chunk_size seconds (default: 30 s) before being sent to Whisper. Two VAD backends are available: pyannote (default) and silero. See Voice Activity Detection for details.
2

Stage 2 — ASR with MLX Whisper

Each merged VAD chunk is transcribed independently by mlx_whisper.transcribe(). The MLXWhisperPipeline class iterates over the chunks, extracts the corresponding audio slice, calls the model, and collects the resulting text and timing into a list of SingleSegment objects.MLX Whisper runs on Apple Silicon’s GPU automatically. The device parameter you pass to load_model() controls only where the VAD model runs — not the Whisper inference itself. Parameters like compute_type, device_index, and threads are accepted for API compatibility but have no effect on MLX inference.The result of this stage is a TranscriptionResult:
3

Stage 3 — Alignment (optional)

Word-level timestamps are not produced by Whisper directly. The alignment stage uses a wav2vec2 model (loaded via load_align_model()) to perform forced alignment, mapping each word in the transcription to its precise position in the audio.Alignment is language-specific. A default wav2vec2 model is selected automatically for 30+ languages. If no model exists for a language, alignment is skipped and you will see an error. Translation output (--task translate) cannot be aligned and skips this stage automatically.The result is an AlignedTranscriptionResult:
4

Stage 4 — Diarization (optional)

When --diarize is enabled, the DiarizationPipeline class uses pyannote to identify who is speaking at each moment in the audio. The default model is pyannote/speaker-diarization-community-1.Speaker labels are then assigned to each word and segment by assign_word_speakers(). This function builds an IntervalTree from the diarization output — a sorted array with binary search — giving O(log n) speaker lookup instead of O(n) linear scan. For long-form content such as multi-hour podcasts, this provides a roughly 228x speedup over a naive scan.Each segment and word in the output gains a "speaker" field (e.g., "SPEAKER_00", "SPEAKER_01").

Memory management

Transcription, alignment, and diarization each load their own model weights. To avoid keeping multiple large models in memory simultaneously, transcribe_task() explicitly deletes each model and calls gc.collect() before loading the next:

Data flow summary