End-to-end speech recognition
End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.
The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.
Related terms
Frequently asked
What is End-to-end speech recognition?
End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.
Why does End-to-end speech recognition matter?
The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.
What terms are related to End-to-end speech recognition?
Closely related concepts: Automatic speech recognition (ASR), Acoustic model, Language model (in ASR) — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →