How it works
After transcription, segments have start and end times for each chunk of speech but no per-word timestamps. Alignment runs a separate wav2vec2 model over the same audio and uses the CTC forced alignment algorithm to map each character to a point in time, then groups characters into words. The result adds awords list to every segment and a flat word_segments list to the top-level result.
Step-by-step guide
1
Transcribe the audio
Start with a standard transcription result:
2
Load the alignment model
Load the wav2vec2 model for the detected language. Pass
result["language"] to use the auto-detected language:load_align_model selects the appropriate wav2vec2 checkpoint automatically based on the language code. You can override it by passing model_name:3
Run alignment
Pass the transcription segments, alignment model, metadata, and audio path to
whispermlx.align:4
Access word timestamps
The aligned result contains per-word timestamps on each segment and in a flat list:
Complete example
Supported languages
Alignment models are available for over 30 languages. The model is selected automatically from the language code:
If no default model exists for a language, pass a custom wav2vec2 model via
--align_model (CLI) or the model_name parameter.
AlignedTranscriptionResult structure
whispermlx.align() returns an AlignedTranscriptionResult:
SingleAlignedSegment extends the base segment with word data:
SingleWordSegment contains:
start, end, and score are optional on SingleWordSegment — they may be absent for words that could not be aligned (e.g., words containing only characters not present in the model vocabulary). Use .get("start") rather than direct key access.Interpolation methods
When a word cannot be aligned directly, whisper(ml)x uses an interpolation strategy to fill in the missing timestamp. Control this with--interpolate_method (CLI) or the interpolate_method parameter:
CLI options
Skipping alignment
--max_line_width, --max_line_count, --highlight_words) are unavailable.
Character-level alignments
chars field to each segment in the JSON output containing per-character timestamps.
Caveats
Next steps
Speaker diarization
Assign speaker labels to aligned word segments.
Output formats
Use word timestamps to generate highlighted SRT/VTT subtitles.