Glossary → Transcription & speech recognition
Transcription & speech recognition
Timestamping
Timestamping marks transcript text with the time it occurs in the recording — per segment, per speaker turn, or per word — so text and audio can be navigated together.
Coarse timestamps (every paragraph or speaker turn) are enough for citation and review; word-level timing unlocks interactive transcripts, precise clips and caption files. Timestamp accuracy comes from the model or a forced-alignment pass.
Related terms
Word-level timestampsWord-level timestamps attach a start and end time to every individual word in a transcript, rather than to who…
Forced alignmentForced alignment is the process of taking an existing transcript and an audio recording and computing exactly …
Subtitle timingSubtitle timing (spotting) is deciding when each caption appears and disappears — synchronized to speech onset…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →