Voice activity detection (VAD)
Voice activity detection (VAD) is the technique of finding which parts of an audio signal contain speech and which are silence, music or noise, so downstream systems process only the spoken segments.
Transcription pipelines use VAD to skip silence (faster, cheaper) and to prevent a known failure mode where speech models hallucinate text during long non-speech stretches. Voice assistants use it to know when you started and stopped talking.
Aggressive VAD thresholds can clip quiet speakers or trailing words; permissive ones waste compute and invite hallucination — production systems tune this trade-off carefully.
Related terms
Frequently asked
What is Voice activity detection?
Voice activity detection (VAD) is the technique of finding which parts of an audio signal contain speech and which are silence, music or noise, so downstream systems process only the spoken segments.
Why does Voice activity detection matter?
Transcription pipelines use VAD to skip silence (faster, cheaper) and to prevent a known failure mode where speech models hallucinate text during long non-speech stretches. Voice assistants use it to know when you started and stopped talking.
What terms are related to Voice activity detection?
Closely related concepts: Hallucination (in transcription), Automatic speech recognition (ASR), Wake word detection — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →