Glossary → Transcription & speech recognition
Transcription & speech recognition
Speaker identification
Speaker identification determines which known person is speaking by matching voices against enrolled voice profiles — unlike diarization, which only separates anonymous speakers.
Identification requires reference audio per person and raises consent and privacy obligations that plain diarization does not. Most transcript workflows get by with diarization plus manually naming 'Speaker 1' once.
Related terms
Speaker diarizationSpeaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker …
Speaker embeddingA speaker embedding is a compact numeric vector representing the characteristics of a voice, such that the sam…
Speaker labelsSpeaker labels are the per-turn tags in a transcript — 'Speaker 1', 'Interviewer', or a real name — produced b…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →