Upload an audio file or paste a link and download a .vtt: every phrase with its start and end time, in the format web players read for captions and interactive transcripts.
How it works
MP3, WAV, M4A, FLAC, OGG, WMA and AAC go in as they are. Phone voice memos and dictation formats work too.
The language is detected for you. Most files finish in under a minute; an hour of audio takes a few.
Choose .vtt in the download menu. SRT, TXT, DOCX and JSON come from the same transcript.
Why audio needs a VTT
People ask for a VTT from audio for four reasons:
<audio> and <video> elements both take a <track>, and players built on them highlight the current line as it plays..vtt beside each recording.VTT as a transcript format
A VTT is plain text you can open in any editor: a WEBVTT first line, then blocks of a time range and the words. If what you really want is the words without times, download TXT instead, or strip an existing file with VTT to TXT. To go the other way for an editor, VTT to SRT converts in the browser.
Audio quality decides how much checking the file needs. A close microphone in a quiet room gives clean text; a phone on a table in a busy room gives more errors, especially on names. Dual-channel call recordings are mixed to one before transcription.
Sample output
WEBVTT 00:00:34.680 --> 00:00:46.500 fourscore and seven years ago our fathers brought forth upon this continent a new nation conceived in in liberty, and dedicated to the proposition 00:00:46.500 --> 00:00:50.300 that all men are created equal. 00:00:50.300 --> 00:00:57.500 Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived 00:00:57.500 --> 00:01:01.620 and so dedicated, can long endure. 00:01:01.620 --> 00:01:07.920 We are met on a great battlefield of that war we have come to dedicate a portion of that field 00:01:07.920 --> 00:01:14.240 as a final resting place for those who here gave their lives that their nation might live 00:01:15.360 --> 00:01:18.800 it is altogether fitting and proper that we should do this
Real output, unedited: a public-domain LibriVox reading of the Gettysburg Address (2 min 41 s) transcribed here in October 2026. The repeated word and the missing capitals are what came back; you would fix those on the transcript before downloading.
Before you rely on it
Each cue is a phrase from the transcript with its start and end time. Most run a few seconds. A passage spoken without a pause can come out as one long cue of up to about half a minute, so split those in your editor if your platform wants two short lines at a time.
Cue times are taken from the recording itself, so the file lines up with the video or audio you uploaded. If you trim the start of the video afterwards, shift the cues by the same amount in your editor.
On public test recordings, about 7 words in 100 differed from the reference for clean English read speech, between 4 and 9 in 100 for Spanish, Portuguese, Russian, French and German, and about 21 in 100 for Hindi. Noisy rooms, music and crosstalk do worse. Method and every number: the benchmark.
Open the transcript, correct a misheard name or number on its line, and the SRT and VTT you download carry the correction with the timing unchanged.
You get the subtitle file. Fonts, colours, position and captions drawn into the picture are done in your video editor or player, which reads this file.
Up to 10 hours or 5 GB per recording. The box on this page takes files up to 500 MB; larger ones continue on the upload page. 99+ languages, detected automatically.
Pricing
The preview is free. When the full transcript needs a pack, one-time packs are $4 for 500 minutes, $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes. Credits never expire, which comes to about $0.29 to $0.48 per hour of audio.
| Recording | $4 or $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 5-minute voice memo | $0.04 | $0.03 | $0.02 |
| 45-minute interview | $0.36 | $0.27 | $0.22 |
| 2-hour recording | $0.96 | $0.72 | $0.58 |
Figures are the share of a pack each recording uses. A button opens secure checkout for that pack. Every export format is included; full details on the pricing page.
FAQ
Upload the MP3 in the box on this page and download .vtt from the transcript when it is ready. The file is created by transcribing the speech; it is not a container conversion.
MP3, WAV, M4A, FLAC, OGG, AAC, WMA and most other audio and video formats, up to 10 hours or 5 GB per file.
Yes. The HTML audio element accepts a track element just like video, and many web audio players use the VTT to show a synchronised transcript.
It is a transcript with a time range on every phrase. Download TXT or DOCX from the same job if you want the words without timing.
No. The cues carry the words and their timing. Speakers are identified and named in the summary, key quotes and chapters.
Every page in this set uses the same upload box and the same export. All subtitle pages
Free preview. VTT, SRT and plain text from the same upload; packs from $4.