Glossary → Transcription & speech recognition
Transcription & speech recognition

Speaker diarization

Speaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker identity so a transcript can label each turn, even without knowing the speakers' names.

Diarization answers 'Speaker 1 vs Speaker 2', not 'Alice vs Bob' — attaching real names is a separate step called speaker identification. Errors concentrate where humans also struggle: overlapping speech, very short interjections and similar-sounding voices.

Recording quality drives diarization more than any algorithm choice: separate microphone tracks per speaker make labels near-perfect, while a single phone in the middle of a meeting is the hard case.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →