Streaming speech recognition
Streaming speech recognition transcribes audio as it arrives, emitting words within moments of them being spoken, as opposed to batch transcription which processes a complete file.
Streaming powers live captions, dictation and voice agents. It trades some accuracy for latency — the model cannot use future context it has not heard yet — and often revises its last few words as more audio arrives.
Batch processing of the same audio is typically more accurate, which is why 'live notes now, clean transcript after' is a common product pattern.
Related terms
Frequently asked
What is Streaming speech recognition?
Streaming speech recognition transcribes audio as it arrives, emitting words within moments of them being spoken, as opposed to batch transcription which processes a complete file.
Why does Streaming speech recognition matter?
Streaming powers live captions, dictation and voice agents. It trades some accuracy for latency — the model cannot use future context it has not heard yet — and often revises its last few words as more audio arrives.
What terms are related to Streaming speech recognition?
Closely related concepts: Batch transcription, Real-time factor (RTF), Live captioning — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →