Skip to main content
whisper(ml)x is a Python library and CLI that brings Whisper-quality transcription to Apple Silicon at native speed. It replaces the standard Whisper backend with mlx-whisper, running directly on the Apple Silicon GPU via unified memory — no CUDA required. Beyond raw transcription, whisper(ml)x adds word-level timestamps through wav2vec2 forced alignment and speaker diarization via pyannote-audio, giving you a complete speech analysis pipeline in a single tool.

Quick Start

Transcribe your first audio file in under five minutes.

Installation

Install with pip or uv and get set up on Apple Silicon.

CLI Reference

Explore every command-line flag and option.

Python API

Integrate transcription, alignment, and diarization into your code.

How it works

whisper(ml)x runs a three-stage pipeline on your audio:
1

VAD — Voice Activity Detection

Audio is preprocessed using pyannote or silero VAD to detect speech segments, skipping silence and reducing inference time.
2

ASR — Transcription

Each speech segment is transcribed using mlx-whisper, which runs natively on the Apple Silicon GPU via MLX’s unified memory architecture.
3

Alignment + Diarization (optional)

Word-level timestamps are computed using wav2vec2 forced alignment. Speaker labels are assigned via pyannote-audio diarization.

Key features

Apple Silicon native

MLX inference runs on the M-series GPU automatically — no device configuration needed for transcription.

Word-level timestamps

Forced alignment via wav2vec2 gives you precise start and end times for every word.

Speaker diarization

Identify who spoke when using pyannote-audio’s speaker diarization pipeline.

30+ languages

Alignment models are available for over 30 languages, from English and Japanese to Arabic and Indonesian.

CLI and Python API

Use whispermlx from the command line or import it directly into your Python application.

Multiple output formats

Export transcripts as SRT, VTT, TXT, TSV, JSON, or all formats at once.

Explore the docs

Pipeline concepts

Understand the VAD → ASR → Alignment → Diarization architecture.

Model reference

Learn about available Whisper models and how short names are resolved.

Transcription guide

Transcribe audio files using both the CLI and Python API.

Diarization guide

Add speaker labels to your transcripts with pyannote-audio.