WebVTT carries more scaffolding than SubRip, and all of it has to go before the words read as prose. This page lists what to take out, and the easier answer when the video still exists.
What is in the file
A caption file is written for a player, so reading it as a document means deleting everything the player needed.
WEBVTT, sometimes with text after it.NOTE is for the file's editor, never shown to a viewer.STYLE block holds appearance rules and can run to many lines.Then the part no deleting solves: cues are cut to a comfortable reading speed, so sentences arrive in halves. On live video, rolling captions repeat each line.
WEBVTT <- delete NOTE exported by the platform <- delete intro <- delete 00:00:09.000 --> 00:00:12.480 line:90% align:center <- delete Welcome back to the second module.
Being straight about it
The browser subtitle tool turns WebVTT into SubRip, since it reads and writes both. It will not give you a .txt: plain text is not among its outputs. For words with nothing around them, it is a text editor or a fresh transcript.
The same limit applies from the other format: see SRT to TXT. To move between caption formats: VTT to SRT, SRT to VTT, the converter overview.
From the source instead
One pass over the original beats repairing a file never meant to be read.
Upload the file or paste a YouTube, podcast host or Apple Podcasts link. Spotify, Kick VOD, Podbean and Zoom are refused straight away.
Continuous paragraphs, word-level timestamps, speaker labels in the transcript view, no rolling duplication.
Plain text for reuse, a real .docx for Word, Markdown for notes. SRT and VTT come from the same job.
A module caption track becomes handouts or revision notes: transcribe an online course.
Long recordings become summaries without anyone rewatching. See transcribe video.
If the video also needs captions, the caption generator produces them from the same upload.
Pricing
Whatever you export, the cost is the same: recording length counts, not file count. a pack from $4 opens the full transcript in full at any length; further minutes come from packs at $8 for 1,000, $12 for 2,000 or $24 for 5,000.
FAQ
Because WebVTT is allowed to carry more. On top of timecodes there can be a header line, comment blocks, style blocks, named cue identifiers, positioning settings appended to the timecode, and inline markup inside the caption text itself. All of it has to come out.
They are cue settings, which tell the player where to draw the caption using values for line, position, alignment and size. They sit on the same line as the timecode, so deleting that whole line removes them with it.
No, not for plain text. WebVTT permits inline markup for emphasis, timing within a cue and voice spans naming who is talking. Strip the tags and keep the words between them, or a text file ends up peppered with fragments no reader wants.
It cannot. The browser subtitle tool converts between SubRip and WebVTT, and those two are the only outputs it offers. Getting to plain text means either editing the file yourself or exporting TXT from a transcript of the original recording.
That is rolling captions, common in live or streamed video, where each cue redraws the previous line plus the new words so the text scrolls. Flattened into a file it reads as heavy duplication, and removing it by hand is slow.
Transcribe the video itself rather than salvaging its captions. Paste a link or upload the file, then export TXT, DOCX or MD. The output is paragraphs, with no duplication from rolling cues and no markup to strip.
Only if you keep them deliberately while editing. Nothing is added for you: a TXT export contains the words alone, speaker labels are shown in the transcript view on the page, and JSON is the only download format that records them.
Up to 10 hours of running time or 5 GB per file. a pack from $4 opens the full transcript, with credits covering anything longer.
Paste a link or drop the file, and read it as a document in minutes.