Cue numbers, timecodes and the gaps between them all have to go, and the fragments left behind need rejoining before they read like sentences. Below is the manual method, plus the shorter path: transcribe the recording once and take TXT, DOCX or MD straight out.
Straight answer first
It converts between caption formats, and plain text is not one of them.
The subtitle tool on this site reads and writes SubRip and WebVTT. Its output menu holds those two options only, so nothing in it yields a .txt, and it seemed better to say so than to send you hunting.
That leaves two routes that work, and the choice depends on whether the original recording still exists.
The manual method
Any editor with regular expression search will do.
--> catches every one.1 00:00:02,000 --> 00:00:04,300 The council voted on the 2 00:00:04,300 --> 00:00:06,100 proposal this morning.
The council voted on the proposal this morning. rejoined by hand: The council voted on the proposal this morning.
The shorter path
Prose written as prose does not need reassembling.
Upload the file, or paste a podcast host, YouTube or Apple Podcasts link. Spotify, Kick VOD, Podbean and Zoom cannot be read and say so at once.
The transcript opens as paragraphs with word-level timestamps; speaker labels show here while you read.
TXT for plain words, DOCX as a real Word file, MD for notes, JSON for structure. SRT and VTT come from the same job.
Checking a quote is far easier in running text, and timestamps in the text let you jump back to the moment.
Take both. The caption generator covers captions, and SRT vs VTT explains which file to hand over.
Pricing
Formats are not priced separately, so a job giving you TXT gives DOCX, SRT, VTT, JSON and MD too. a pack from $4 opens the full transcript in full at any length; after that, packs of 1,000 minutes at $8, 2,000 at $12 or 5,000 at $24, with no expiry.
FAQ
Three kinds of line. The counter above each cue, the timecode line containing the arrow, and the blank line separating one cue from the next. What remains is the spoken words, still broken at the points a caption had to break.
No, and it is worth saying plainly. That tool reads and writes caption formats only, so its output menu offers SubRip and WebVTT and nothing else. For plain text, either export TXT from a transcript or strip the file by hand in an editor.
Open the .srt in any editor with regular expression search. Delete lines matching a pattern for the arrow timecode, delete lines that contain only digits, then collapse the repeated blank lines. It takes about a minute and needs no installation.
Because captions are broken for reading speed on screen, not for grammar. A sentence is often split across two or three cues, so stripping the timings leaves short fragments that need rejoining before the text reads as prose.
Yes, when you still have the audio or video. Transcribe the recording and export TXT, DOCX or MD directly. The text is written as continuous paragraphs rather than reassembled from caption fragments, so nothing needs repairing afterwards.
No. A TXT download carries the words alone, and an SRT never held speaker information to begin with. Speaker labels are shown in the transcript view on the page, and of the download formats only JSON records them.
Usually reading, searching or reusing. A word count for a script, a quote checked against what was actually said, a summary drafted from a talk, or the text pasted into a document where timecodes would only be noise.
A pack from $4 opens the full transcript. Files can run to 10 hours or 5 GB, and anything past that first free transcript draws on credits which do not expire.
Drop the recording or paste a link, and take TXT, DOCX and MD from the same transcript.