load_align_model, then running align.
load_align_model
(model, metadata) tuple to pass directly to align().
Parameters
string
required
BCP-47 language code for the audio (e.g.
"en", "fr", "ja"). A default wav2vec2 model is selected automatically for 30+ languages including English, French, German, Spanish, Italian, Japanese, Chinese, Arabic, Hindi, and more.string
required
PyTorch device for the alignment model. Use
"cpu" or "mps".string
Override the default alignment model for the language. Pass any wav2vec2 model name from Hugging Face or a torchaudio pipeline name (e.g.
"WAV2VEC2_ASR_BASE_960H").string
Directory to cache downloaded alignment model files. Defaults to the Hugging Face cache directory.
boolean
default:"false"
When
True, load alignment models from local cache only — no network requests. Raises an error if the model is not already cached. Useful for offline environments.Returns
A tuple(model, metadata):
torch.nn.Module
required
The loaded wav2vec2 model, moved to
device. Pass this directly to align().dict
required
Dictionary with keys
"language", "dictionary", and "type". Pass this directly to align().Full signature
align
AlignedTranscriptionResult with per-word start and end times.
Parameters
list[SingleSegment]
required
The
segments list from a TranscriptionResult (i.e. result["segments"]).torch.nn.Module
required
Alignment model returned by
load_align_model.dict
required
Metadata dict returned by
load_align_model.str | np.ndarray | torch.Tensor
required
The same audio used for transcription. A file path string, a 16 kHz mono float32 numpy array, or a PyTorch tensor. File paths are loaded automatically via
load_audio.string
required
PyTorch device for alignment inference (
"cpu" or "mps"). Should match the device used in load_align_model.boolean
default:"false"
When
True, each aligned segment includes a chars field with per-character timestamps in addition to word timestamps.string
default:"nearest"
Method used to fill in timestamps for words that could not be directly aligned.
"nearest" assigns the timestamp of the nearest aligned neighbor.boolean
default:"false"
Print percentage progress to stdout as each segment is aligned.
callable
Called with a float between
0.0 and 100.0 after each segment is processed. See ProgressCallback.Returns
AnAlignedTranscriptionResult TypedDict:
list[SingleAlignedSegment]
required
Aligned segments. Each segment contains a
words list with per-word timestamps.list[SingleWordSegment]
required
Flat list of all word segments across every aligned segment. Convenient for word-level iteration without nesting.
Full signature
Usage
When
task="translate" is used during transcription, the output text is in English but the audio is not — forced alignment cannot be performed on translated output. Skip alignment in this case.Related
MLXWhisperPipeline.transcribe
Produce the transcript that alignment operates on.
Diarization
Assign speaker labels to aligned segments.
Schema & Types
Full TypedDict definitions for AlignedTranscriptionResult.