TranscriptionResult with segment-level text and timestamps.
Parameters
str | np.ndarray
required
Audio input. Either:
- A file path string (any format supported by ffmpeg: mp3, wav, m4a, flac, ogg, etc.)
- A NumPy array of shape
(samples,)withdtype=float32at 16 kHz mono. Useload_audioto prepare a file as a numpy array.
string
BCP-47 language code override (e.g.
"en", "fr"). When None, the language is auto-detected from the first processed chunk. Overrides the language set on the pipeline at load time.string
"transcribe" or "translate". Overrides the task set on the pipeline at load time. When set to "translate", output is always in English.integer
default:"30"
Maximum audio chunk size in seconds used when merging VAD segments. Smaller values reduce peak memory usage at the cost of more model calls.
boolean
default:"false"
Print percentage progress to stdout after each VAD chunk is processed.
boolean
default:"false"
Print each segment’s text and timestamps to stdout as transcription progresses.
callable
A callable with signature
(pct: float) -> None. Called after each VAD segment is processed with a value between 0.0 and 100.0. See ProgressCallback.Returns
ATranscriptionResult TypedDict:
list[SingleSegment]
required
One entry per VAD chunk.
string
required
The language code used for transcription (e.g.
"en"). Equal to the detected language when language was not specified.Full signature
Usage
To get word-level timestamps, pass
result["segments"] to whispermlx.align() after transcription.Related
load_model
Load a model to get an MLXWhisperPipeline.
Alignment
Add word-level timestamps to transcription output.
Schema & Types
Full TypedDict definitions for TranscriptionResult.
Audio Utilities
Load audio files as numpy arrays.