Glossary → Transcription & speech recognition
Transcription & speech recognition
End-to-end speech recognition
End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.
The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.
Related terms
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…
Acoustic modelAn acoustic model is the component of a speech-recognition system that maps audio features to the sounds of sp…
Language model (in ASR)In speech recognition, a language model supplies knowledge of which word sequences are likely, letting the sys…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →