speaker-diarization-community-1 model to perform this step.
Prerequisites
The diarization model is gated on Hugging Face. Before using it, complete these steps:1
Create a Hugging Face account and generate an access token
Go to huggingface.co/settings/tokens and create a token with at least read access.
2
Accept the model agreement
Visit the pyannote/speaker-diarization-community-1 model page on Hugging Face and accept the usage agreement. You must be signed in.
3
Pass the token to whisper(ml)x
Provide your token via
--hf_token on the CLI or the token parameter on DiarizationPipeline.CLI usage
Enable diarization with the--diarize flag and supply your token:
Python API
The full three-stage pipeline — transcribe, align, diarize — gives the best results. Alignment is recommended before diarization because word-level timestamps allow speaker assignment at word granularity.Speaker count hints
If you know the number of speakers in advance, passmin_speakers and/or max_speakers to improve accuracy:
num_speakers if you know the exact count:
Speaker embeddings
Request speaker embedding vectors alongside the diarization result usingreturn_embeddings=True. The embeddings dictionary maps each speaker ID to a list of floats representing that speaker’s voice profile:
--speaker_embeddings:
Progress tracking
Pass aprogress_callback to DiarizationPipeline.__call__() to receive progress updates. The callback maps two internal pyannote steps — segmentation and embedding — into a monotonic 0–100 range:
How speaker assignment works
assign_word_speakers maps each diarization segment (a time interval with a speaker label) onto transcription segments and individual words.
For every segment or word, it queries which diarization intervals overlap and picks the speaker with the greatest total overlap duration. If a segment has no overlapping diarization interval and fill_nearest=True is set, it falls back to the nearest segment by midpoint.
Internally,
assign_word_speakers uses a custom IntervalTree backed by a sorted array and binary search. This gives O(log n) query time per word, which is approximately 228× faster than a linear scan — a meaningful improvement for long recordings such as multi-hour podcasts or meetings.DiarizationPipeline API
assign_word_speakers API
Next steps
Output formats
Write speaker-labeled transcripts to SRT, JSON, and other formats.
Alignment guide
Learn about word-level alignment, which enables per-word speaker labels.