Glossary → Transcription & speech recognition
Transcription & speech recognition
Forced alignment
Forced alignment is the process of taking an existing transcript and an audio recording and computing exactly when each word (or phoneme) was spoken, producing precise word-level timestamps.
Unlike plain recognition, alignment already knows the words — the task is only timing. That is why tools like WhisperX layer a dedicated alignment model on top of Whisper: Whisper's own timestamps are approximate, and alignment sharpens them to tens of milliseconds.
Accurate alignment is what makes karaoke-style captions, clickable transcripts that seek the audio, precise clipping and dubbing workflows possible.
Related terms
Word-level timestampsWord-level timestamps attach a start and end time to every individual word in a transcript, rather than to who…
PhonemeA phoneme is the smallest sound unit that distinguishes words in a language — English has roughly 44, spanning…
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →