Glossary → Transcription & speech recognition
Transcription & speech recognition

Automatic speech recognition (ASR)

Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-readable text, powering transcription, captions, dictation and voice assistants.

Modern ASR systems are end-to-end neural networks — models like OpenAI's Whisper are trained on hundreds of thousands of hours of multilingual audio and output punctuated text directly, replacing the older pipeline of separate acoustic and language models.

ASR quality varies with audio clarity, accent coverage, vocabulary and language resources: a clean podcast in English transcribes near-perfectly, while noisy phone audio in a low-resource language still deserves a human review pass.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →