Glossary → Transcription & speech recognition
Transcription & speech recognition
Acoustic model
An acoustic model is the component of a speech-recognition system that maps audio features to the sounds of speech, historically paired with a separate language model that mapped sounds to likely word sequences.
The acoustic/language model split defined classical ASR for decades. Modern end-to-end systems fold both roles into one neural network, but the vocabulary survives — 'acoustic' issues still mean audio-side problems (noise, accents, microphone) as opposed to vocabulary-side ones.
Related terms
Language model (in ASR)In speech recognition, a language model supplies knowledge of which word sequences are likely, letting the sys…
End-to-end speech recognitionEnd-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the cla…
Mel spectrogramA mel spectrogram is the time-frequency picture of audio that speech models actually consume — a spectrogram w…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →