Glossary → Transcription & speech recognition
Transcription & speech recognition
Word-level timestamps
Word-level timestamps attach a start and end time to every individual word in a transcript, rather than to whole sentences or segments.
Per-word timing enables interactive transcripts (click a word, the audio seeks there), tight caption timing, silence-trimmed clips and edit-by-text video workflows.
Recognition models emit approximate word times; when tighter precision is needed, a forced-alignment pass refines them against the audio.
Related terms
Forced alignmentForced alignment is the process of taking an existing transcript and an audio recording and computing exactly …
TimestampingTimestamping marks transcript text with the time it occurs in the recording — per segment, per speaker turn, o…
Subtitle timingSubtitle timing (spotting) is deciding when each caption appears and disappears — synchronized to speech onset…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →