Glossary → Transcription & speech recognition
Transcription & speech recognition

End-to-end speech recognition

End-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the classical pipeline of separate acoustic, pronunciation and language models.

The end-to-end approach — transformer models like Whisper are the best-known examples — simplified ASR dramatically and made massive multilingual training practical. Its trade-off is opacity: there is no separate component to patch when a specific word keeps failing, which is why vocabulary boosting and post-processing still exist.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →