Skip to main content
Speaker diarization answers the question “who spoke when?” The DiarizationPipeline class uses pyannote to segment audio by speaker, and assign_word_speakers merges those speaker labels into a transcript.

DiarizationPipeline

Wraps a pyannote speaker diarization model. Instantiate once, then call the instance on audio files.

Constructor

string
default:"pyannote/speaker-diarization-community-1"
Pyannote diarization model name on Hugging Face. Defaults to "pyannote/speaker-diarization-community-1".
string
Hugging Face access token. Required for gated models (including the default). Generate one at huggingface.co/settings/tokens and accept the model’s terms of use.
string | torch.device
default:"mps"
Device for diarization inference. Use "cpu" or "mps". Defaults to "mps" for Apple Silicon.
string
Directory to cache downloaded model files. Defaults to the Hugging Face cache directory.

Calling the pipeline

str | np.ndarray
required
Audio input. Either a file path string or a 16 kHz mono float32 numpy array.
integer
Exact number of speakers, if known. When provided, the model skips speaker count estimation.
integer
Minimum number of speakers to detect. Used when the exact count is unknown.
integer
Maximum number of speakers to detect. Used when the exact count is unknown.
boolean
default:"false"
When True, the pipeline returns a tuple (diarize_df, speaker_embeddings) instead of just diarize_df. The embeddings dict maps speaker IDs to float lists.
callable
Called with a float between 0.0 and 100.0 as diarization progresses. Progress is split across two internal stages: segmentation (0–50%) and embeddings (50–99%). See ProgressCallback.

Returns

When return_embeddings=False (default): a pandas DataFrame with columns: When return_embeddings=True: a tuple (DataFrame, dict[str, list[float]]). The dict maps each speaker ID to its embedding vector.

Full signature


assign_word_speakers

Merges speaker labels from the diarization DataFrame into a transcript result. Uses an interval tree for O(log n) speaker lookup — significantly faster than a linear scan for long recordings. Each segment and word in the transcript gains a "speaker" field (e.g. "SPEAKER_00").

Parameters

pd.DataFrame
required
Diarization DataFrame returned by DiarizationPipeline.
AlignedTranscriptionResult | TranscriptionResult
required
The transcript to augment. Typically the result of align(), but a plain TranscriptionResult from transcribe() is also accepted (speaker labels will be added to segments only, not words).
dict[str, list[float]]
Optional speaker embeddings dict from DiarizationPipeline (when return_embeddings=True). When provided, the embeddings are stored on the result under the key "speaker_embeddings".
boolean
default:"false"
When True, assign the nearest speaker even when there is no direct time overlap between a word/segment and any diarization segment. Useful for audio with brief pauses between speaker segments.

Returns

The transcript_result dict augmented in-place and returned. Each segment dict gains a "speaker" key, and each word dict (in aligned transcripts) also gains a "speaker" key.

Full signature


Usage

The default pyannote diarization model (pyannote/speaker-diarization-community-1) is gated. You must accept its terms on Hugging Face and pass a valid token to DiarizationPipeline. Without a token, model loading will fail.

Alignment

Align transcription to word-level timestamps before diarization.

Schema & Types

Full TypedDict definitions for AlignedTranscriptionResult.

Diarization guide

End-to-end walkthrough of the full pipeline.