DiarizationPipeline class uses pyannote to segment audio by speaker, and assign_word_speakers merges those speaker labels into a transcript.
DiarizationPipeline
Constructor
string
default:"pyannote/speaker-diarization-community-1"
Pyannote diarization model name on Hugging Face. Defaults to
"pyannote/speaker-diarization-community-1".string
Hugging Face access token. Required for gated models (including the default). Generate one at huggingface.co/settings/tokens and accept the model’s terms of use.
string | torch.device
default:"mps"
Device for diarization inference. Use
"cpu" or "mps". Defaults to "mps" for Apple Silicon.string
Directory to cache downloaded model files. Defaults to the Hugging Face cache directory.
Calling the pipeline
str | np.ndarray
required
Audio input. Either a file path string or a 16 kHz mono float32 numpy array.
integer
Exact number of speakers, if known. When provided, the model skips speaker count estimation.
integer
Minimum number of speakers to detect. Used when the exact count is unknown.
integer
Maximum number of speakers to detect. Used when the exact count is unknown.
boolean
default:"false"
When
True, the pipeline returns a tuple (diarize_df, speaker_embeddings) instead of just diarize_df. The embeddings dict maps speaker IDs to float lists.callable
Called with a float between
0.0 and 100.0 as diarization progresses. Progress is split across two internal stages: segmentation (0–50%) and embeddings (50–99%). See ProgressCallback.Returns
Whenreturn_embeddings=False (default): a pandas DataFrame with columns:
When
return_embeddings=True: a tuple (DataFrame, dict[str, list[float]]). The dict maps each speaker ID to its embedding vector.
Full signature
assign_word_speakers
"speaker" field (e.g. "SPEAKER_00").
Parameters
pd.DataFrame
required
Diarization DataFrame returned by
DiarizationPipeline.AlignedTranscriptionResult | TranscriptionResult
required
The transcript to augment. Typically the result of
align(), but a plain TranscriptionResult from transcribe() is also accepted (speaker labels will be added to segments only, not words).dict[str, list[float]]
Optional speaker embeddings dict from
DiarizationPipeline (when return_embeddings=True). When provided, the embeddings are stored on the result under the key "speaker_embeddings".boolean
default:"false"
When
True, assign the nearest speaker even when there is no direct time overlap between a word/segment and any diarization segment. Useful for audio with brief pauses between speaker segments.Returns
Thetranscript_result dict augmented in-place and returned. Each segment dict gains a "speaker" key, and each word dict (in aligned transcripts) also gains a "speaker" key.
Full signature
Usage
Related
Alignment
Align transcription to word-level timestamps before diarization.
Schema & Types
Full TypedDict definitions for AlignedTranscriptionResult.
Diarization guide
End-to-end walkthrough of the full pipeline.