Glossary → Transcription & speech recognition
Transcription & speech recognition
Real-time factor (RTF)
Real-time factor (RTF) measures transcription speed as processing time divided by audio duration: an RTF of 0.1 means one hour of audio transcribes in six minutes; values below 1 are faster than real time.
Speed claims are often quoted the other way up ('30x real time' means RTF ≈ 0.033). RTF depends on hardware, model size and settings — the same model can be 50x real time on a GPU and slower than real time on a laptop CPU.
For live captioning, RTF must stay well under 1 continuously; for batch transcription, RTF just sets how long you wait.
Related terms
Streaming speech recognitionStreaming speech recognition transcribes audio as it arrives, emitting words within moments of them being spok…
Batch transcriptionBatch transcription processes complete, already-recorded audio files, letting the system use full context in b…
Quantization (of speech models)Quantization shrinks a neural speech model by storing its weights at lower numeric precision — for example INT…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →