Skip to main content
Transcribes audio by running VAD to segment speech, then passing each chunk through MLX Whisper. Returns a TranscriptionResult with segment-level text and timestamps.

Parameters

str | np.ndarray
required
Audio input. Either:
  • A file path string (any format supported by ffmpeg: mp3, wav, m4a, flac, ogg, etc.)
  • A NumPy array of shape (samples,) with dtype=float32 at 16 kHz mono. Use load_audio to prepare a file as a numpy array.
string
BCP-47 language code override (e.g. "en", "fr"). When None, the language is auto-detected from the first processed chunk. Overrides the language set on the pipeline at load time.
string
"transcribe" or "translate". Overrides the task set on the pipeline at load time. When set to "translate", output is always in English.
integer
default:"30"
Maximum audio chunk size in seconds used when merging VAD segments. Smaller values reduce peak memory usage at the cost of more model calls.
boolean
default:"false"
Print percentage progress to stdout after each VAD chunk is processed.
boolean
default:"false"
Print each segment’s text and timestamps to stdout as transcription progresses.
callable
A callable with signature (pct: float) -> None. Called after each VAD segment is processed with a value between 0.0 and 100.0. See ProgressCallback.

Returns

A TranscriptionResult TypedDict:
list[SingleSegment]
required
One entry per VAD chunk.
string
required
The language code used for transcription (e.g. "en"). Equal to the detected language when language was not specified.

Full signature

Usage

To get word-level timestamps, pass result["segments"] to whispermlx.align() after transcription.

load_model

Load a model to get an MLXWhisperPipeline.

Alignment

Add word-level timestamps to transcription output.

Schema & Types

Full TypedDict definitions for TranscriptionResult.

Audio Utilities

Load audio files as numpy arrays.