whispermlx.schema. Import them for type annotations or runtime isinstance checks.
TranscriptionResult
Returned byMLXWhisperPipeline.transcribe().
list[SingleSegment]
required
One segment per VAD chunk. Each segment covers a continuous span of speech.
string
required
The language code used for transcription (e.g.
"en", "fr"). Equal to the auto-detected language when no language was specified at load or call time.AlignedTranscriptionResult
Returned bywhispermlx.align().
list[SingleAlignedSegment]
required
Segments with word-level timestamps added. See
SingleAlignedSegment.list[SingleWordSegment]
required
Flat list of all word segments across every segment. Useful for iterating words without nesting.
SingleSegment
One segment of speech produced bytranscribe().
float
required
Segment start time in seconds.
float
required
Segment end time in seconds.
string
required
Transcribed text for this segment.
float
Average log probability across all tokens in this segment. More negative values indicate lower model confidence. Not always present.
SingleAlignedSegment
A segment after forced alignment — extendsSingleSegment with word and character timestamps.
float
required
Segment start time in seconds.
float
required
Segment end time in seconds.
string
required
Segment text.
list[SingleWordSegment]
required
Word-level timestamp entries for this segment. See
SingleWordSegment.list[SingleCharSegment] | None
Per-character timestamp entries. Only populated when
return_char_alignments=True is passed to align(). None otherwise. See SingleCharSegment.float
Average log probability from Whisper. Not always present.
SingleWordSegment
A single word with alignment timestamps.string
required
The word text, including surrounding punctuation if any.
float
Word start time in seconds. May be absent if the word could not be aligned.
float
Word end time in seconds. May be absent if the word could not be aligned.
float
Alignment confidence score between 0 and 1. Higher is more confident.
SingleCharSegment
A single character with alignment timestamps. Only present whenreturn_char_alignments=True is passed to align().
string
required
The character.
float
Character start time in seconds.
float
Character end time in seconds.
float
Alignment confidence score between 0 and 1.
ProgressCallback
Type alias for progress callbacks accepted bytranscribe(), align(), and DiarizationPipeline.
float argument in the range 0.0 to 100.0 representing the percentage of work completed. Passing None disables progress reporting.
Usage examples
Logging
whisper(ml)x uses a centralized logger under the"whispermlx" namespace. Two functions are exported from the top-level package for configuring and accessing it.
setup_logging
"whispermlx" logger. Call this once at the start of your application before any other whisper(ml)x calls.
string
default:"info"
Logging level. Accepted values:
"debug", "info", "warning", "error", "critical".string
Optional path to a log file. When set, log output is written to both the console and this file. When
None, logs only to stdout.get_logger
logging.Logger instance configured with the whisper(ml)x handler. Use this to emit log messages that follow the same format as internal whisper(ml)x logs.
string
required
Logger name. Typically pass
__name__ from your calling module.The logger propagation is disabled to prevent duplicate messages in environments that configure the root logger. If you use a custom logging setup, call
setup_logging() after configuring the root logger.Related
MLXWhisperPipeline.transcribe
Returns TranscriptionResult.
Alignment
Returns AlignedTranscriptionResult.
Diarization
Adds speaker fields to segments and words.