Word error rate (WER)
Word error rate (WER) is the standard accuracy metric for speech recognition: the number of word substitutions, deletions and insertions needed to turn the system's transcript into the correct one, divided by the number of words actually spoken.
A WER of 5% means roughly one error every twenty words. Because WER counts every divergence — including harmless ones like 'okay' vs 'OK' — a system can post a low WER and still read badly, or post a mediocre WER while capturing every fact correctly.
WER is only comparable when measured on the same audio: vendors quoting WER on clean read speech will always beat numbers measured on real meetings, accents and background noise.
Related terms
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →