Glossary → Transcription & speech recognition
Transcription & speech recognition
Transcription
Transcription is the process of converting spoken audio into written text, either by a human typist or by automatic speech recognition software.
Modern transcription is dominated by neural speech-recognition models that convert hours of audio in minutes, with human review reserved for the highest-stakes contexts like certified legal transcripts.
The practical quality bar is whether the text is accurate enough for its purpose: searchable notes tolerate small errors, published quotes and captions do not — which is why word error rate and a quick review pass both matter.
Related terms
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…
Word error rate (WER)Word error rate (WER) is the standard accuracy metric for speech recognition: the number of word substitutions…
Verbatim transcriptionVerbatim transcription captures speech exactly as uttered — including false starts, repetitions, filler words …
Speaker diarizationSpeaker diarization is the process of determining who spoke when in a recording — segmenting audio by speaker …
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →