Glossary → Transcription & speech recognition
Transcription & speech recognition

Speaker labels

Speaker labels are the per-turn tags in a transcript — 'Speaker 1', 'Interviewer', or a real name — produced by diarization and optionally renamed by a human or an identification step.

Good tools let you rename a label once and apply it everywhere. Label errors cluster in cross-talk and short interjections, so a quick scan of turn boundaries is the highest-value review step for multi-speaker audio.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →