Glossary → Transcription & speech recognition
Transcription & speech recognition

Forced alignment

Forced alignment is the process of taking an existing transcript and an audio recording and computing exactly when each word (or phoneme) was spoken, producing precise word-level timestamps.

Unlike plain recognition, alignment already knows the words — the task is only timing. That is why tools like WhisperX layer a dedicated alignment model on top of Whisper: Whisper's own timestamps are approximate, and alignment sharpens them to tens of milliseconds.

Accurate alignment is what makes karaoke-style captions, clickable transcripts that seek the audio, precise clipping and dubbing workflows possible.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →