Glossary → Transcription & speech recognition
Transcription & speech recognition

Multilingual speech recognition

Multilingual speech recognition uses one model trained across many languages — Whisper's published list covers 99 — so a single system can transcribe, and often translate, audio in most of the world's major languages.

Accuracy is not uniform: high-resource languages (English, Spanish, German) perform best, while low-resource languages are supported but deserve a review pass. Multilingual training is also what makes code-switching manageable.

The practical win is operational: one pipeline for a global archive instead of per-language vendors.

Related terms

Put the term to work

Transcribe audio or video in 99 languages — speaker labels, word timestamps, captions.

Transcribe a file →