Glossary → Transcription & speech recognition
Transcription & speech recognition

Transcription

Transcription is the process of converting spoken audio into written text, either by a human typist or by automatic speech recognition software.

Modern transcription is dominated by neural speech-recognition models that convert hours of audio in minutes, with human review reserved for the highest-stakes contexts like certified legal transcripts.

The practical quality bar is whether the text is accurate enough for its purpose: searchable notes tolerate small errors, published quotes and captions do not — which is why word error rate and a quick review pass both matter.

Related terms

Frequently asked

What is Transcription?

Transcription is the process of converting spoken audio into written text, either by a human typist or by automatic speech recognition software.

Why does Transcription matter?

Modern transcription is dominated by neural speech-recognition models that convert hours of audio in minutes, with human review reserved for the highest-stakes contexts like certified legal transcripts.

What terms are related to Transcription?

Closely related concepts: Automatic speech recognition (ASR), Word error rate (WER), Verbatim transcription, Speaker diarization — each has its own entry in this glossary.

Put the term to work

Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.

Transcribe a file →