Glossary → Transcription & speech recognition
Transcription & speech recognition

Word-level timestamps

Word-level timestamps attach a start and end time to every individual word in a transcript, rather than to whole sentences or segments.

Per-word timing enables interactive transcripts (click a word, the audio seeks there), tight caption timing, silence-trimmed clips and edit-by-text video workflows.

Recognition models emit approximate word times; when tighter precision is needed, a forced-alignment pass refines them against the audio.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →