Glossary → Transcription & speech recognition
Transcription & speech recognition

Acoustic model

An acoustic model is the component of a speech-recognition system that maps audio features to the sounds of speech, historically paired with a separate language model that mapped sounds to likely word sequences.

The acoustic/language model split defined classical ASR for decades. Modern end-to-end systems fold both roles into one neural network, but the vocabulary survives — 'acoustic' issues still mean audio-side problems (noise, accents, microphone) as opposed to vocabulary-side ones.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →