Glossary → Transcription & speech recognition
Transcription & speech recognition
Voice activity detection (VAD)
Voice activity detection (VAD) is the technique of finding which parts of an audio signal contain speech and which are silence, music or noise, so downstream systems process only the spoken segments.
Transcription pipelines use VAD to skip silence (faster, cheaper) and to prevent a known failure mode where speech models hallucinate text during long non-speech stretches. Voice assistants use it to know when you started and stopped talking.
Aggressive VAD thresholds can clip quiet speakers or trailing words; permissive ones waste compute and invite hallucination — production systems tune this trade-off carefully.
Related terms
Hallucination (in transcription)In speech recognition, hallucination is when the model outputs fluent text that was never spoken — typically r…
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…
Wake word detectionWake word detection is the tiny, always-on model that listens for a trigger phrase ('Hey Siri') and only then …
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →