Glossary → Transcription & speech recognition
Transcription & speech recognition

Forced alignment

Forced alignment is the process of taking an existing transcript and an audio recording and computing exactly when each word (or phoneme) was spoken, producing precise word-level timestamps.

Unlike plain recognition, alignment already knows the words — the task is only timing. That is why tools like WhisperX layer a dedicated alignment model on top of Whisper: Whisper's own timestamps are approximate, and alignment sharpens them to tens of milliseconds.

Accurate alignment is what makes karaoke-style captions, clickable transcripts that seek the audio, precise clipping and dubbing workflows possible.

Related terms

Frequently asked

What is Forced alignment?

Forced alignment is the process of taking an existing transcript and an audio recording and computing exactly when each word (or phoneme) was spoken, producing precise word-level timestamps.

Why does Forced alignment matter?

Unlike plain recognition, alignment already knows the words — the task is only timing. That is why tools like WhisperX layer a dedicated alignment model on top of Whisper: Whisper's own timestamps are approximate, and alignment sharpens them to tens of milliseconds.

What terms are related to Forced alignment?

Closely related concepts: Word-level timestamps, Phoneme, Automatic speech recognition (ASR) — each has its own entry in this glossary.

Put the term to work

Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.

Transcribe a file →