How often speech models invent text
A hallucination study from production transcripts: what the artifacts are, how often they appear, and in which languages.
What this measures
Speech models sometimes write text that was never said: a phrase repeated dozens of times over silence, subtitle credits learned from training data, a run of filler where nothing was spoken. We run a deterministic detector over every transcript and then a second check that decides whether to keep, trim or remove each flagged span. This page reports the detector over the last 90 days of finished transcripts (2,851 of them), by the language the engine detected. Method on how we test; the detector's thresholds are the ones in production.
Since the cleaning pass went live, of 61 transcripts it examined, 32.8 % had at least one segment removed or trimmed after the second check.
By language
"Flagged transcripts" had at least one suspicious span; "flagged segments" is the share of segments the detector marked. Languages with fewer than ten transcripts in the window are not shown.
| Language | Transcripts | Audio hours | Flagged transcripts | Flagged segments |
|---|---|---|---|---|
| English | 1,368 | 1376.1 | 27.7 % | 0.22 % |
| Filipino / Tagalog | 60 | 33.6 | 68.3 % | 0.89 % |
| Spanish | 45 | 51.0 | 42.2 % | 0.22 % |
| Hindi | 31 | 6.8 | 41.9 % | 2.63 % |
| Portuguese | 28 | 22.0 | 25.0 % | 0.09 % |
| Russian | 28 | 23.1 | 25.0 % | 0.08 % |
| Chinese | 27 | 21.7 | 51.9 % | 0.27 % |
| French | 25 | 19.8 | 24.0 % | 0.1 % |
| German | 17 | 13.0 | 11.8 % | 0.09 % |
| Japanese | 15 | 7.1 | 60.0 % | 0.4 % |
| Norwegian (Nynorsk, as detected) | 12 | 5.9 | 58.3 % | 4.98 % |
| Tamil | 12 | 2.8 | 66.7 % | 0.76 % |
| Arabic | 10 | 3.9 | 70.0 % | 1.93 % |
| az | 10 | 1.5 | 20.0 % | 0.4 % |
What the artifacts look like
| Artifact | Spans flagged |
|---|---|
| The same segment twice in a row Two consecutive segments with identical text; usually a decoding hiccup at a chunk boundary. | 2,355 |
| A phrase repeated many times The model gets stuck on a phrase and emits it eight, twenty, fifty times; happens on silence, music beds and long pauses. | 709 |
| Subtitle credits that were never spoken 'Subtitles by …', 'Thanks for watching' and their equivalents in other languages, learned from subtitle training data. | 220 |
| A segment in the wrong script A single segment in a script neither neighbour uses. | 49 |
| A long run with almost no distinct words Twenty or more tokens with under a quarter distinct; filler where there was no speech. | 38 |
Examples, as removed
Real spans the second check removed, shortened; the original transcript is kept and the removed span shows in the detailed view.
So the earth was full, blah,→ [removed] (English, repeat_loop)but, it's all good.→ [removed] (English, repeat_loop)So the earth was full, blah,→ [removed] (English, repeat_loop)but, it's all good.→ [removed] (English, repeat_loop)कुछ अभी चेंजेस हुए हैं तो वह आपको थोड़ा सपोर्ट करें डॉक्यूमेंट यह तो यह सब जी सर अब अमें म→ [removed] (Hindi, repeat_loop,low_diversity)काम करते हैं तो रोड पंटेनर जहां पर राम होता है वहां गार्डन वेस्ट ऐसे अलग-अलग उनके और जो हम→ [removed] (Hindi, repeat_loop)
Why we publish it
An accuracy percentage without its failure modes is not useful. These are the failure modes, with their frequency, from production, refreshed monthly. Quote them with the date.
Related measurements
Accuracy and speed on public test sets: benchmark. What people send in a month: what people transcribe. Method: how we test.
See also: Transcription services · Turnaround time · Evidence log · Pricing · Privacy · Terms