Glossary → Transcription & speech recognition
Transcription & speech recognition

Voice activity detection (VAD)

Voice activity detection (VAD) is the technique of finding which parts of an audio signal contain speech and which are silence, music or noise, so downstream systems process only the spoken segments.

Transcription pipelines use VAD to skip silence (faster, cheaper) and to prevent a known failure mode where speech models hallucinate text during long non-speech stretches. Voice assistants use it to know when you started and stopped talking.

Aggressive VAD thresholds can clip quiet speakers or trailing words; permissive ones waste compute and invite hallucination — production systems tune this trade-off carefully.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →