Glossary → Transcription & speech recognition
Transcription & speech recognition

End-to-end speech recognition

End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.

The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.

Related terms

Frequently asked

What is End-to-end speech recognition?

End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.

Why does End-to-end speech recognition matter?

The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.

What terms are related to End-to-end speech recognition?

Closely related concepts: Automatic speech recognition (ASR), Acoustic model, Language model (in ASR) — each has its own entry in this glossary.

Put the term to work

Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.

Transcribe a file →