Skip to main content
Speaker diarization answers the question “who spoke when?” It assigns a speaker label to each segment and word in a transcript. whisper(ml)x uses pyannote-audio’s speaker-diarization-community-1 model to perform this step.

Prerequisites

The diarization model is gated on Hugging Face. Before using it, complete these steps:
1

Create a Hugging Face account and generate an access token

Go to huggingface.co/settings/tokens and create a token with at least read access.
2

Accept the model agreement

Visit the pyannote/speaker-diarization-community-1 model page on Hugging Face and accept the usage agreement. You must be signed in.
3

Pass the token to whisper(ml)x

Provide your token via --hf_token on the CLI or the token parameter on DiarizationPipeline.

CLI usage

Enable diarization with the --diarize flag and supply your token:
Provide speaker count hints if you know how many speakers are present:

Python API

The full three-stage pipeline — transcribe, align, diarize — gives the best results. Alignment is recommended before diarization because word-level timestamps allow speaker assignment at word granularity.

Speaker count hints

If you know the number of speakers in advance, pass min_speakers and/or max_speakers to improve accuracy:
Pass num_speakers if you know the exact count:

Speaker embeddings

Request speaker embedding vectors alongside the diarization result using return_embeddings=True. The embeddings dictionary maps each speaker ID to a list of floats representing that speaker’s voice profile:
To include embeddings in the JSON output file via the CLI, add --speaker_embeddings:

Progress tracking

Pass a progress_callback to DiarizationPipeline.__call__() to receive progress updates. The callback maps two internal pyannote steps — segmentation and embedding — into a monotonic 0–100 range:

How speaker assignment works

assign_word_speakers maps each diarization segment (a time interval with a speaker label) onto transcription segments and individual words. For every segment or word, it queries which diarization intervals overlap and picks the speaker with the greatest total overlap duration. If a segment has no overlapping diarization interval and fill_nearest=True is set, it falls back to the nearest segment by midpoint.
Internally, assign_word_speakers uses a custom IntervalTree backed by a sorted array and binary search. This gives O(log n) query time per word, which is approximately 228× faster than a linear scan — a meaningful improvement for long recordings such as multi-hour podcasts or meetings.

DiarizationPipeline API

assign_word_speakers API

Next steps

Output formats

Write speaker-labeled transcripts to SRT, JSON, and other formats.

Alignment guide

Learn about word-level alignment, which enables per-word speaker labels.