Glossary → Transcription & speech recognition
Transcription & speech recognition
Multilingual speech recognition
Multilingual speech recognition uses one model trained across many languages — Whisper's published list covers 99 — so a single system can transcribe, and often translate, audio in most of the world's major languages.
Accuracy is not uniform: high-resource languages (English, Spanish, German) perform best, while low-resource languages are supported but deserve a review pass. Multilingual training is also what makes code-switching manageable.
The practical win is operational: one pipeline for a global archive instead of per-language vendors.
Related terms
Code-switchingCode-switching is when a speaker alternates between two or more languages within a conversation or a single se…
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…
Word error rate (WER)Word error rate (WER) is the standard accuracy metric for speech recognition: the number of word substitutions…
Put the term to work
Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.
Transcribe a file →