Glossary → Transcription & speech recognition
Transcription & speech recognition
Streaming speech recognition
Streaming speech recognition transcribes audio as it arrives, emitting words within moments of them being spoken, as opposed to batch transcription which processes a complete file.
Streaming powers live captions, dictation and voice agents. It trades some accuracy for latency — the model cannot use future context it has not heard yet — and often revises its last few words as more audio arrives.
Batch processing of the same audio is typically more accurate, which is why 'live notes now, clean transcript after' is a common product pattern.
Related terms
Batch transcriptionBatch transcription processes complete, already-recorded audio files, letting the system use full context in b…
Real-time factor (RTF)Real-time factor (RTF) measures transcription speed as processing time divided by audio duration: an RTF of 0.…
Live captioningLive captioning produces captions in real time for broadcasts, streams, meetings and events, via streaming spe…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →