Glossary → Voice & AI audio
Voice & AI audio
Audio classification
Audio classification labels what a sound is — speech, music, applause, glass breaking — without transcribing any words.
Classifiers tag content for search, trigger events in monitoring, and route pipelines (send speech to ASR, skip music). Caption workflows use them to describe non-speech sounds.
Related terms
Voice activity detection (VAD)Voice activity detection (VAD) is the technique of finding which parts of an audio signal contain speech and w…
Source separationSource separation splits mixed audio into components — vocals from music, one speaker from another — using neu…
Frequently asked
What is Audio classification?
Audio classification labels what a sound is — speech, music, applause, glass breaking — without transcribing any words.
Why does Audio classification matter?
Classifiers tag content for search, trigger events in monitoring, and route pipelines (send speech to ASR, skip music). Caption workflows use them to describe non-speech sounds.
What terms are related to Audio classification?
Closely related concepts: Voice activity detection (VAD), Source separation — each has its own entry in this glossary.