Glossary → Transcription & speech recognition
Transcription & speech recognition
Language model (in ASR)
In speech recognition, a language model supplies knowledge of which word sequences are likely, letting the system prefer 'recognize speech' over 'wreck a nice beach' when the audio alone is ambiguous.
Classical systems used explicit n-gram language models; end-to-end systems learn the same knowledge implicitly. Domain-specific boosts (medical terms, product names) are the modern descendant — telling the recognizer which unusual words to expect.
Related terms
Acoustic modelAn acoustic model is the component of a speech-recognition system that maps audio features to the sounds of sp…
Beam searchBeam search is a decoding strategy where a speech-recognition model keeps several candidate transcriptions in …
End-to-end speech recognitionEnd-to-end speech recognition uses a single neural network that maps audio directly to text, replacing the cla…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →