A VTT (WebVTT) file is the caption format web browsers read. It is plain text: a WEBVTT line, then time ranges and words. Make one from a recording in the box below, or read on for the layout and the rules.
How it works
Drop a video or audio file into the box, or paste a link.
Fix anything misheard on the transcript; timing is measured from the audio.
Choose .vtt in the download menu and host the file beside your video.
Anatomy
WEBVTT. Nothing may come before it. A blank line follows.-->, end, written hours:minutes:seconds.milliseconds with a full stop before the milliseconds. The hours may be left out for times under an hour.The format can do more than SRT: cue settings after the time range place text on screen (line:10% align:start), NOTE blocks hold comments, and STYLE blocks hold CSS. Files made here are the simple, widely compatible kind: header and timed text, no settings. You add appearance in your page's CSS with ::cue.
Serve the file as text/vtt, saved as UTF-8.
Checking a VTT
WEBVTT on line one, browsers ignore the entire file.00:00:05,000 is SRT style; a VTT parser drops that cue.--> inside cue text is forbidden, since it marks a time range.To test, add the file to a video with a track element and open the page through a web server; the browser's console reports parsing errors. Markup and the common loading failures are on subtitles for HTML5 video.
Converting between the two formats loses nothing for plain cues: VTT to SRT, SRT to VTT. The differences in one table: SRT vs VTT. The SRT side of this explainer: the SRT file format.
Sample output
WEBVTT 00:00:34.680 --> 00:00:46.500 fourscore and seven years ago our fathers brought forth upon this continent a new nation conceived in in liberty, and dedicated to the proposition 00:00:46.500 --> 00:00:50.300 that all men are created equal. 00:00:50.300 --> 00:00:57.500 Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived 00:00:57.500 --> 00:01:01.620 and so dedicated, can long endure. 00:01:01.620 --> 00:01:07.920 We are met on a great battlefield of that war we have come to dedicate a portion of that field 00:01:07.920 --> 00:01:14.240 as a final resting place for those who here gave their lives that their nation might live 00:01:15.360 --> 00:01:18.800 it is altogether fitting and proper that we should do this
Real output, unedited: a public-domain LibriVox reading of the Gettysburg Address (2 min 41 s) transcribed here in October 2026. The repeated word and the missing capitals are what came back; you would fix those on the transcript before downloading.
Before you rely on it
Each cue is a phrase from the transcript with its start and end time. Most run a few seconds. A passage spoken without a pause can come out as one long cue of up to about half a minute, so split those in your editor if your platform wants two short lines at a time.
Cue times are taken from the recording itself, so the file lines up with the video or audio you uploaded. If you trim the start of the video afterwards, shift the cues by the same amount in your editor.
Open the transcript, correct a misheard name or number on its line, and the SRT and VTT you download carry the correction with the timing unchanged.
You get the subtitle file. Fonts, colours, position and captions drawn into the picture are done in your video editor or player, which reads this file.
On public test recordings, about 7 words in 100 differed from the reference for clean English read speech, between 4 and 9 in 100 for Spanish, Portuguese, Russian, French and German, and about 21 in 100 for Hindi. Noisy rooms, music and crosstalk do worse. Method and every number: the benchmark.
Up to 10 hours or 5 GB per recording. The box on this page takes files up to 500 MB; larger ones continue on the upload page. 99+ languages, detected automatically.
Pricing
The preview is free. When the full transcript needs a pack, one-time packs are $4 for 500 minutes, $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes. Credits never expire, which comes to about $0.29 to $0.48 per hour of audio.
| Recording | $4 or $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 5-minute video | $0.04 | $0.03 | $0.02 |
| 30-minute video | $0.24 | $0.18 | $0.14 |
| 2-hour recording | $0.96 | $0.72 | $0.58 |
Figures are the share of a pack each recording uses. A button opens secure checkout for that pack. Every export format is included; full details on the pricing page.
FAQ
A VTT file is a WebVTT caption file: plain text that begins with the line WEBVTT and lists timed cues. It is the format web browsers use to show captions and subtitles on HTML5 video.
Yes. WebVTT files carry subtitles, captions, chapters or descriptions for web video. Platforms such as Vimeo and Udemy accept them as caption uploads.
Upload your video or audio in the box on this page and download the .vtt made from its speech. You can also write one by hand: the WEBVTT line, a blank line, then time ranges and text.
Any text editor opens it for reading and editing. To see it with a video, load it in a web page with a track element, or in a player such as VLC.
VTT requires a WEBVTT header, uses a full stop before milliseconds, and supports positioning and styling. SRT numbers its cues, uses a comma, and has no header. The timed text itself is the same.
Every page in this set uses the same upload box and the same export. All subtitle pages
Free preview. Drop a file, download the .vtt; packs from $4.