Glossary → Transcription & speech recognition
Transcription & speech recognition

Automatic speech recognition (ASR)

Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-readable text, powering transcription, captions, dictation and voice assistants.

Modern ASR systems are end-to-end neural networks — models like OpenAI's Whisper are trained on hundreds of thousands of hours of multilingual audio and output punctuated text directly, replacing the older pipeline of separate acoustic and language models.

ASR quality varies with audio clarity, accent coverage, vocabulary and language resources: a clean podcast in English transcribes near-perfectly, while noisy phone audio in a low-resource language still deserves a human review pass.

Related terms

Frequently asked

What is Automatic speech recognition?

Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-readable text, powering transcription, captions, dictation and voice assistants.

Why does Automatic speech recognition matter?

Modern ASR systems are end-to-end neural networks — models like OpenAI's Whisper are trained on hundreds of thousands of hours of multilingual audio and output punctuated text directly, replacing the older pipeline of separate acoustic and language models.

What terms are related to Automatic speech recognition?

Closely related concepts: Word error rate (WER), Acoustic model, Language model (in ASR), End-to-end speech recognition, Streaming speech recognition — each has its own entry in this glossary.

Put the term to work

Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.

Transcribe a file →