Audio to VTT: MP3, WAV or M4A to a WebVTT file.

Upload an audio file or paste a link and download a .vtt: every phrase with its start and end time, in the format web players read for captions and interactive transcripts.

MP3, WAV, M4A, FLAC, OGG.vtt and .srtUp to 10 hoursFree preview

How it works

From an audio file to a VTT in three steps.

1

Upload the audio

MP3, WAV, M4A, FLAC, OGG, WMA and AAC go in as they are. Phone voice memos and dictation formats work too.

2

Let it transcribe

The language is detected for you. Most files finish in under a minute; an hour of audio takes a few.

3

Download .vtt

Choose .vtt in the download menu. SRT, TXT, DOCX and JSON come from the same transcript.

Why audio needs a VTT

A timed text track is not only for video.

People ask for a VTT from audio for four reasons:

VTT as a transcript format

Reading it, converting it, trimming it.

A VTT is plain text you can open in any editor: a WEBVTT first line, then blocks of a time range and the words. If what you really want is the words without times, download TXT instead, or strip an existing file with VTT to TXT. To go the other way for an editor, VTT to SRT converts in the browser.

Audio quality decides how much checking the file needs. A close microphone in a quiet room gives clean text; a phone on a table in a busy room gives more errors, especially on names. Dual-channel call recordings are mixed to one before transcription.

Sample output

What the VTT file looks like.

WEBVTT

00:00:34.680 --> 00:00:46.500
fourscore and seven years ago our fathers brought forth upon this continent a new nation conceived in in liberty, and dedicated to the proposition

00:00:46.500 --> 00:00:50.300
that all men are created equal.

00:00:50.300 --> 00:00:57.500
Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived

00:00:57.500 --> 00:01:01.620
and so dedicated, can long endure.

00:01:01.620 --> 00:01:07.920
We are met on a great battlefield of that war we have come to dedicate a portion of that field

00:01:07.920 --> 00:01:14.240
as a final resting place for those who here gave their lives that their nation might live

00:01:15.360 --> 00:01:18.800
it is altogether fitting and proper that we should do this

Real output, unedited: a public-domain LibriVox reading of the Gettysburg Address (2 min 41 s) transcribed here in October 2026. The repeated word and the missing capitals are what came back; you would fix those on the transcript before downloading.

Before you rely on it

What the subtitle file is, and what it is not.

One cue per spoken phrase

Each cue is a phrase from the transcript with its start and end time. Most run a few seconds. A passage spoken without a pause can come out as one long cue of up to about half a minute, so split those in your editor if your platform wants two short lines at a time.

Timing comes from the audio

Cue times are taken from the recording itself, so the file lines up with the video or audio you uploaded. If you trim the start of the video afterwards, shift the cues by the same amount in your editor.

Accuracy we measured

On public test recordings, about 7 words in 100 differed from the reference for clean English read speech, between 4 and 9 in 100 for Spanish, Portuguese, Russian, French and German, and about 21 in 100 for Hindi. Noisy rooms, music and crosstalk do worse. Method and every number: the benchmark.

Fix a line before you download

Open the transcript, correct a misheard name or number on its line, and the SRT and VTT you download carry the correction with the timing unchanged.

A plain file, not a styled video

You get the subtitle file. Fonts, colours, position and captions drawn into the picture are done in your video editor or player, which reads this file.

Limits

Up to 10 hours or 5 GB per recording. The box on this page takes files up to 500 MB; larger ones continue on the upload page. 99+ languages, detected automatically.

Pricing

What a VTT file from audio costs

The preview is free. When the full transcript needs a pack, one-time packs are $4 for 500 minutes, $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes. Credits never expire, which comes to about $0.29 to $0.48 per hour of audio.

Recording$4 or $8 pack$12 pack$24 pack
5-minute voice memo$0.04$0.03$0.02
45-minute interview$0.36$0.27$0.22
2-hour recording$0.96$0.72$0.58

Figures are the share of a pack each recording uses. A button opens secure checkout for that pack. Every export format is included; full details on the pricing page.

FAQ

Audio to VTT, answered.

How do I convert an MP3 to a VTT file?

Upload the MP3 in the box on this page and download .vtt from the transcript when it is ready. The file is created by transcribing the speech; it is not a container conversion.

Which audio formats can I upload?

MP3, WAV, M4A, FLAC, OGG, AAC, WMA and most other audio and video formats, up to 10 hours or 5 GB per file.

Can a VTT be used with an audio-only player?

Yes. The HTML audio element accepts a track element just like video, and many web audio players use the VTT to show a synchronised transcript.

Is the VTT the same as a transcript?

It is a transcript with a time range on every phrase. Download TXT or DOCX from the same job if you want the words without timing.

Does it label speakers in the VTT?

No. The cues carry the words and their timing. Speakers are identified and named in the summary, key quotes and chapters.

Subtitles and captions

Every page in this set uses the same upload box and the same export. All subtitle pages

Closest to this page

WebVTT

Make subtitles

From a link or a file

For a platform or an editor

Convert a subtitle file

Formats and rules

By language

Drop the audio. Get the VTT.

Free preview. VTT, SRT and plain text from the same upload; packs from $4.