Skip to main content
All public types are defined in whispermlx.schema. Import them for type annotations or runtime isinstance checks.

TranscriptionResult

Returned by MLXWhisperPipeline.transcribe().
list[SingleSegment]
required
One segment per VAD chunk. Each segment covers a continuous span of speech.
string
required
The language code used for transcription (e.g. "en", "fr"). Equal to the auto-detected language when no language was specified at load or call time.

AlignedTranscriptionResult

Returned by whispermlx.align().
list[SingleAlignedSegment]
required
Segments with word-level timestamps added. See SingleAlignedSegment.
list[SingleWordSegment]
required
Flat list of all word segments across every segment. Useful for iterating words without nesting.

SingleSegment

One segment of speech produced by transcribe().
float
required
Segment start time in seconds.
float
required
Segment end time in seconds.
string
required
Transcribed text for this segment.
float
Average log probability across all tokens in this segment. More negative values indicate lower model confidence. Not always present.

SingleAlignedSegment

A segment after forced alignment — extends SingleSegment with word and character timestamps.
float
required
Segment start time in seconds.
float
required
Segment end time in seconds.
string
required
Segment text.
list[SingleWordSegment]
required
Word-level timestamp entries for this segment. See SingleWordSegment.
list[SingleCharSegment] | None
Per-character timestamp entries. Only populated when return_char_alignments=True is passed to align(). None otherwise. See SingleCharSegment.
float
Average log probability from Whisper. Not always present.

SingleWordSegment

A single word with alignment timestamps.
string
required
The word text, including surrounding punctuation if any.
float
Word start time in seconds. May be absent if the word could not be aligned.
float
Word end time in seconds. May be absent if the word could not be aligned.
float
Alignment confidence score between 0 and 1. Higher is more confident.

SingleCharSegment

A single character with alignment timestamps. Only present when return_char_alignments=True is passed to align().
string
required
The character.
float
Character start time in seconds.
float
Character end time in seconds.
float
Alignment confidence score between 0 and 1.

ProgressCallback

Type alias for progress callbacks accepted by transcribe(), align(), and DiarizationPipeline.
The callable receives a single float argument in the range 0.0 to 100.0 representing the percentage of work completed. Passing None disables progress reporting.

Usage examples


Logging

whisper(ml)x uses a centralized logger under the "whispermlx" namespace. Two functions are exported from the top-level package for configuring and accessing it.

setup_logging

Configures the "whispermlx" logger. Call this once at the start of your application before any other whisper(ml)x calls.
string
default:"info"
Logging level. Accepted values: "debug", "info", "warning", "error", "critical".
string
Optional path to a log file. When set, log output is written to both the console and this file. When None, logs only to stdout.

get_logger

Returns a logging.Logger instance configured with the whisper(ml)x handler. Use this to emit log messages that follow the same format as internal whisper(ml)x logs.
string
required
Logger name. Typically pass __name__ from your calling module.
The logger propagation is disabled to prevent duplicate messages in environments that configure the root logger. If you use a custom logging setup, call setup_logging() after configuring the root logger.

MLXWhisperPipeline.transcribe

Returns TranscriptionResult.

Alignment

Returns AlignedTranscriptionResult.

Diarization

Adds speaker fields to segments and words.