Glossary → Voice & AI audio
Voice & AI audio
Speaker embedding
A speaker embedding is a compact numeric vector representing the characteristics of a voice, such that the same speaker's clips land close together — the machinery behind diarization and speaker verification.
Embeddings turn 'same voice?' into a distance measurement. Diarization clusters them; verification compares them to an enrolled profile; both inherit their biases from the embedding model's training data.
Related terms
Speaker diarizationSpeaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker …
Speaker identificationSpeaker identification determines which known person is speaking by matching voices against enrolled voice pro…
Voice cloningVoice cloning creates a synthetic voice that imitates a specific real person, sometimes from minutes of refere…