Glossary → Transcription & speech recognition
Transcription & speech recognition

Acoustic model

An acoustic model is the component of a speech-recognition system that maps audio features to the sounds of speech, historically paired with a separate language model that mapped sounds to likely word sequences.

The acoustic/language model split defined classical ASR for decades. Modern end-to-end systems fold both roles into one neural network, but the vocabulary survives — 'acoustic' issues still mean audio-side problems (noise, accents, microphone) as opposed to vocabulary-side ones.

Related terms

Frequently asked

What is Acoustic model?

An acoustic model is the component of a speech-recognition system that maps audio features to the sounds of speech, historically paired with a separate language model that mapped sounds to likely word sequences.

Why does Acoustic model matter?

The acoustic/language model split defined classical ASR for decades. Modern end-to-end systems fold both roles into one neural network, but the vocabulary survives — 'acoustic' issues still mean audio-side problems (noise, accents, microphone) as opposed to vocabulary-side ones.

What terms are related to Acoustic model?

Closely related concepts: Language model (in ASR), End-to-end speech recognition, Mel spectrogram — each has its own entry in this glossary.

Put the term to work

Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.

Transcribe a file →