Glossary → Transcription & speech recognition
Transcription & speech recognition

Streaming speech recognition

Streaming speech recognition transcribes audio as it arrives, emitting words within moments of them being spoken, as opposed to batch transcription which processes a complete file.

Streaming powers live captions, dictation and voice agents. It trades some accuracy for latency — the model cannot use future context it has not heard yet — and often revises its last few words as more audio arrives.

Batch processing of the same audio is typically more accurate, which is why 'live notes now, clean transcript after' is a common product pattern.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →