Glossary → Transcription & speech recognition
Transcription & speech recognition
Speaker diarization
Speaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker identity so a transcript can label each turn, even without knowing the speakers' names.
Diarization answers 'Speaker 1 vs Speaker 2', not 'Alice vs Bob' — attaching real names is a separate step called speaker identification. Errors concentrate where humans also struggle: overlapping speech, very short interjections and similar-sounding voices.
Recording quality drives diarization more than any algorithm choice: separate microphone tracks per speaker make labels near-perfect, while a single phone in the middle of a meeting is the hard case.
Related terms
Speaker identificationSpeaker identification determines which known person is speaking by matching voices against enrolled voice pro…
Speaker labelsSpeaker labels are the per-turn tags in a transcript — 'Speaker 1', 'Interviewer', or a real name — produced b…
TranscriptionTranscription is the process of converting spoken audio into written text, either by a human typist or by auto…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →