Skip to main content
whisper(ml)x is a fork of WhisperX by Max Bain et al. that replaces the inference backend with mlx-whisper, running natively on Apple Silicon via MLX. Word-level timestamps, speaker diarization, and voice activity detection are all retained from the original.

Why use whisper(ml)x

MLX inference

Runs transcription on the Apple Silicon GPU via unified memory — no CUDA, no discrete GPU required.

Word-level timestamps

Forced alignment via wav2vec2 produces precise start and end times for every word in the transcript.

Speaker diarization

Identify who spoke when using pyannote-audio’s speaker diarization pipeline.

VAD preprocessing

Voice activity detection (pyannote or silero) strips silence before transcription, reducing inference time on long recordings.

Architecture overview

whisper(ml)x runs a three-stage pipeline on your audio:
1

VAD — Voice activity detection

Audio is preprocessed by pyannote or silero VAD to detect speech segments. Silence is discarded and speech chunks are merged into 30-second windows before being passed to the ASR stage.
2

ASR — Transcription

Each speech segment is transcribed by mlx-whisper, which runs natively on the Apple Silicon GPU. The result is a list of SingleSegment dicts containing start, end, and text.
3

Alignment and diarization (optional)

A wav2vec2 alignment model attaches word-level timestamps to each segment, producing SingleAlignedSegment dicts. Optionally, a pyannote diarization model assigns speaker labels to each word.

System requirements

  • Apple Silicon Mac — M1, M2, M3, or M4 chip (required for MLX inference)
  • macOS — any version supported by your chip
  • Python — 3.10, 3.11, 3.12, or 3.13
MLX only runs on Apple Silicon. whisper(ml)x will not work on Intel Macs or non-macOS platforms for transcription. The device parameter controls PyTorch models (VAD, alignment, diarization) only — MLX Whisper always uses the Apple Silicon GPU automatically.

Get started

Quick start

Transcribe your first audio file in under five minutes.

Installation

Install with pip or uv and configure speaker diarization.

CLI reference

Explore every command-line flag and option.

Python API

Integrate transcription, alignment, and diarization into your code.

Acknowledgements

whisper(ml)x is built on top of WhisperX by Max Bain et al., mlx-whisper, pyannote-audio, and OpenAI Whisper.