Speaker diarization
Speaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker identity so a transcript can label each turn, even without knowing the speakers' names.
Diarization answers 'Speaker 1 vs Speaker 2', not 'Alice vs Bob' — attaching real names is a separate step called speaker identification. Errors concentrate where humans also struggle: overlapping speech, very short interjections and similar-sounding voices.
Recording quality drives diarization more than any algorithm choice: separate microphone tracks per speaker make labels near-perfect, while a single phone in the middle of a meeting is the hard case.
Related terms
Frequently asked
What is Speaker diarization?
Speaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker identity so a transcript can label each turn, even without knowing the speakers' names.
Why does Speaker diarization matter?
Diarization answers 'Speaker 1 vs Speaker 2', not 'Alice vs Bob' — attaching real names is a separate step called speaker identification. Errors concentrate where humans also struggle: overlapping speech, very short interjections and similar-sounding voices.
What terms are related to Speaker diarization?
Closely related concepts: Speaker identification, Speaker labels, Transcription — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →