Skip to main content
Forced alignment maps Whisper’s segment-level transcriptions to precise word timestamps by running a wav2vec2 model against the original audio. The process requires two steps: loading a language-specific model with load_align_model, then running align.

load_align_model

Loads a wav2vec2 alignment model for the given language. Returns a (model, metadata) tuple to pass directly to align().

Parameters

string
required
BCP-47 language code for the audio (e.g. "en", "fr", "ja"). A default wav2vec2 model is selected automatically for 30+ languages including English, French, German, Spanish, Italian, Japanese, Chinese, Arabic, Hindi, and more.
string
required
PyTorch device for the alignment model. Use "cpu" or "mps".
string
Override the default alignment model for the language. Pass any wav2vec2 model name from Hugging Face or a torchaudio pipeline name (e.g. "WAV2VEC2_ASR_BASE_960H").
string
Directory to cache downloaded alignment model files. Defaults to the Hugging Face cache directory.
boolean
default:"false"
When True, load alignment models from local cache only — no network requests. Raises an error if the model is not already cached. Useful for offline environments.

Returns

A tuple (model, metadata):
torch.nn.Module
required
The loaded wav2vec2 model, moved to device. Pass this directly to align().
dict
required
Dictionary with keys "language", "dictionary", and "type". Pass this directly to align().

Full signature


align

Aligns a list of transcription segments to word-level timestamps using forced alignment. Returns an AlignedTranscriptionResult with per-word start and end times.

Parameters

list[SingleSegment]
required
The segments list from a TranscriptionResult (i.e. result["segments"]).
torch.nn.Module
required
Alignment model returned by load_align_model.
dict
required
Metadata dict returned by load_align_model.
str | np.ndarray | torch.Tensor
required
The same audio used for transcription. A file path string, a 16 kHz mono float32 numpy array, or a PyTorch tensor. File paths are loaded automatically via load_audio.
string
required
PyTorch device for alignment inference ("cpu" or "mps"). Should match the device used in load_align_model.
boolean
default:"false"
When True, each aligned segment includes a chars field with per-character timestamps in addition to word timestamps.
string
default:"nearest"
Method used to fill in timestamps for words that could not be directly aligned. "nearest" assigns the timestamp of the nearest aligned neighbor.
boolean
default:"false"
Print percentage progress to stdout as each segment is aligned.
callable
Called with a float between 0.0 and 100.0 after each segment is processed. See ProgressCallback.

Returns

An AlignedTranscriptionResult TypedDict:
list[SingleAlignedSegment]
required
Aligned segments. Each segment contains a words list with per-word timestamps.
list[SingleWordSegment]
required
Flat list of all word segments across every aligned segment. Convenient for word-level iteration without nesting.

Full signature


Usage

Alignment is language-specific and requires a wav2vec2 model trained for the target language. If no default model exists for a language, load_align_model raises a ValueError. You can pass a custom model_name for unsupported languages.
When task="translate" is used during transcription, the output text is in English but the audio is not — forced alignment cannot be performed on translated output. Skip alignment in this case.

MLXWhisperPipeline.transcribe

Produce the transcript that alignment operates on.

Diarization

Assign speaker labels to aligned segments.

Schema & Types

Full TypedDict definitions for AlignedTranscriptionResult.