Multilingual speech recognition
Multilingual speech recognition uses one model trained across many languages — Whisper's published list covers 99 — so a single system can transcribe, and often translate, audio in most of the world's major languages.
Accuracy is not uniform: high-resource languages (English, Spanish, German) perform best, while low-resource languages are supported but deserve a review pass. Multilingual training is also what makes code-switching manageable.
The practical win is operational: one pipeline for a global archive instead of per-language vendors.
Related terms
Frequently asked
What is Multilingual speech recognition?
Multilingual speech recognition uses one model trained across many languages — Whisper's published list covers 99 — so a single system can transcribe, and often translate, audio in most of the world's major languages.
Why does Multilingual speech recognition matter?
Accuracy is not uniform: high-resource languages (English, Spanish, German) perform best, while low-resource languages are supported but deserve a review pass. Multilingual training is also what makes code-switching manageable.
What terms are related to Multilingual speech recognition?
Closely related concepts: Code-switching, Automatic speech recognition (ASR), Word error rate (WER) — each has its own entry in this glossary.
Transcribe audio or video in 99+ languages — speaker labels, word timestamps, captions.
Transcribe a file →