Quick Start
Transcribe your first audio file in under five minutes.
Installation
Install with pip or uv and get set up on Apple Silicon.
CLI Reference
Explore every command-line flag and option.
Python API
Integrate transcription, alignment, and diarization into your code.
How it works
whisper(ml)x runs a three-stage pipeline on your audio:1
VAD — Voice Activity Detection
Audio is preprocessed using pyannote or silero VAD to detect speech segments, skipping silence and reducing inference time.
2
ASR — Transcription
Each speech segment is transcribed using mlx-whisper, which runs natively on the Apple Silicon GPU via MLX’s unified memory architecture.
3
Alignment + Diarization (optional)
Word-level timestamps are computed using wav2vec2 forced alignment. Speaker labels are assigned via pyannote-audio diarization.
Key features
Apple Silicon native
MLX inference runs on the M-series GPU automatically — no device configuration needed for transcription.
Word-level timestamps
Forced alignment via wav2vec2 gives you precise start and end times for every word.
Speaker diarization
Identify who spoke when using pyannote-audio’s speaker diarization pipeline.
30+ languages
Alignment models are available for over 30 languages, from English and Japanese to Arabic and Indonesian.
CLI and Python API
Use whispermlx from the command line or import it directly into your Python application.
Multiple output formats
Export transcripts as SRT, VTT, TXT, TSV, JSON, or all formats at once.
Explore the docs
Pipeline concepts
Understand the VAD → ASR → Alignment → Diarization architecture.
Model reference
Learn about available Whisper models and how short names are resolved.
Transcription guide
Transcribe audio files using both the CLI and Python API.
Diarization guide
Add speaker labels to your transcripts with pyannote-audio.