Quantization (of speech models)
Quantization shrinks a neural speech model by storing its weights at lower numeric precision — for example INT8 instead of 32-bit floats — cutting memory use and speeding up inference with a small accuracy cost.
Quantization is why large models run on laptops and phones at all: an INT8 Whisper variant needs a fraction of the memory of the original. The accuracy loss is usually small on clear audio and shows up first on hard audio and rare words.
Related terms
Frequently asked
What is Quantization?
Quantization shrinks a neural speech model by storing its weights at lower numeric precision — for example INT8 instead of 32-bit floats — cutting memory use and speeding up inference with a small accuracy cost.
Why does Quantization matter?
Quantization is why large models run on laptops and phones at all: an INT8 Whisper variant needs a fraction of the memory of the original. The accuracy loss is usually small on clear audio and shows up first on hard audio and rare words.
What terms are related to Quantization?
Closely related concepts: Real-time factor (RTF), End-to-end speech recognition — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →