Skip to main content
Alignment takes a transcription result and computes the exact start and end time for every word. It uses a wav2vec2 phoneme recognition model to perform forced alignment between the audio signal and the text, producing per-word timestamps with confidence scores.

How it works

After transcription, segments have start and end times for each chunk of speech but no per-word timestamps. Alignment runs a separate wav2vec2 model over the same audio and uses the CTC forced alignment algorithm to map each character to a point in time, then groups characters into words. The result adds a words list to every segment and a flat word_segments list to the top-level result.

Step-by-step guide

1

Transcribe the audio

Start with a standard transcription result:
2

Load the alignment model

Load the wav2vec2 model for the detected language. Pass result["language"] to use the auto-detected language:
load_align_model selects the appropriate wav2vec2 checkpoint automatically based on the language code. You can override it by passing model_name:
3

Run alignment

Pass the transcription segments, alignment model, metadata, and audio path to whispermlx.align:
4

Access word timestamps

The aligned result contains per-word timestamps on each segment and in a flat list:

Complete example

Supported languages

Alignment models are available for over 30 languages. The model is selected automatically from the language code: If no default model exists for a language, pass a custom wav2vec2 model via --align_model (CLI) or the model_name parameter.

AlignedTranscriptionResult structure

whispermlx.align() returns an AlignedTranscriptionResult:
Each SingleAlignedSegment extends the base segment with word data:
Each SingleWordSegment contains:
start, end, and score are optional on SingleWordSegment — they may be absent for words that could not be aligned (e.g., words containing only characters not present in the model vocabulary). Use .get("start") rather than direct key access.

Interpolation methods

When a word cannot be aligned directly, whisper(ml)x uses an interpolation strategy to fill in the missing timestamp. Control this with --interpolate_method (CLI) or the interpolate_method parameter:

CLI options

Skipping alignment

When alignment is skipped, output formats that depend on word-level timestamps (--max_line_width, --max_line_count, --highlight_words) are unavailable.

Character-level alignments

This adds a chars field to each segment in the JSON output containing per-character timestamps.

Caveats

The translate task is incompatible with alignment. When --task translate is set, alignment is disabled automatically. Forced alignment requires the transcript text to match the spoken words in the original language.
Japanese (ja) and Chinese (zh) do not use spaces between words. For these languages, each character is treated as its own word during alignment. Segment boundaries and word counts will differ from space-delimited languages.

Next steps

Speaker diarization

Assign speaker labels to aligned word segments.

Output formats

Use word timestamps to generate highlighted SRT/VTT subtitles.