Home›How we test

How we test

Windows, definitions and sources for every number on the site, so you can judge them.

Every number on this site has a window and a source

When a page says "an hour of audio came back in 4.1 minutes", that is a measurement over a stated window of production jobs, refreshed automatically, not a figure chosen for a brochure. This page explains how each kind of number is computed so you can judge it.

Turnaround

Definition. Time from the moment the recording has fully landed on our servers to the moment the transcript is written, divided by the recording's length. Queue time, when every GPU is busy, is not included; the upload page shows queue position live, and the turnaround page reports the 90-day view including it.

Reported as the median (half of recordings were faster) and the 90th percentile (nine in ten were faster), over the last 30 days of finished jobs. Currently: 4.1 minutes and 20.1 minutes per hour of audio.

Accuracy

We run a fixed set of reference recordings with known transcripts through the production engine every week and compare word by word. The reference set is private on purpose: publishing it would let anyone optimise against it. What we publish is the shape of the errors, not a single accuracy percentage stripped of its conditions, because accuracy on clean studio audio and accuracy on a phone recording in a café are different numbers. Names, specialist jargon, heavy accents, crosstalk and poor audio are where errors appear; every word is timestamped so you can jump to the moment and listen. For public, reproducible numbers see the benchmark on FLEURS and LibriSpeech (twelve languages, five-minute files, through the customer upload path) and the hallucination study by language.

Since September 2026 a cleaning pass checks each transcript for the artifacts speech models produce on silence and music (a phrase repeated many times, subtitle-credit lines that were never spoken). In the last two weeks it changed something in about a third of transcripts and about a quarter of one percent of segments. The original is kept; removed segments show in the detailed view.

Languages, formats and sources

"Languages detected" counts the languages the engine identified across finished recordings in the window, not a catalogue of languages the model claims to support. Formats are the file types people actually uploaded. Sources are where recordings came from: uploads, the browser recorder, the API, or a link to the person's own cloud storage or a public feed.

Competitor prices and features

Every comparative statement on the site is logged in the evidence log with the page it was taken from and the date it was checked. Prices change; the date tells you how fresh the claim is. If a vendor's page moves or its price changes, the log entry is corrected, not deleted.

What we do not publish