Captions · timings removed

Get the words out of an SRT as plain text.

Cue numbers, timecodes and the gaps between them all have to go, and the fragments left behind need rejoining before they read like sentences. Below is the manual method, plus the shorter path: transcribe the recording once and take TXT, DOCX or MD straight out.

TXT, DOCX, MD exports Real .docx for Word Searchable library 10 hours per file

Straight answer first

Our free browser tool will not do this one.

It converts between caption formats, and plain text is not one of them.

The subtitle tool on this site reads and writes SubRip and WebVTT. Its output menu holds those two options only, so nothing in it yields a .txt, and it seemed better to say so than to send you hunting.

That leaves two routes that work, and the choice depends on whether the original recording still exists.

You only have the caption file. Strip it by hand in a text editor: a one-minute job with regular expression search.
You still have the audio or video. Transcribe it and export TXT. The text arrives as paragraphs, with nothing to repair.

The manual method

Stripping an SRT in a text editor.

Any editor with regular expression search will do.

  1. Delete the timecode lines. A regular expression anchored on --> catches every one.
  2. Delete the counters. Remove lines of digits alone. These cue numbers are why a naive copy and paste comes out sprinkled with stray numerals.
  3. Collapse the blank lines. Replace runs of two or more empty lines with one, which turns the file into continuous text.
  4. Rejoin the sentences. No tool does this well: captions break where the screen needs a break, so join the fragments by eye.

Before

1
00:00:02,000 --> 00:00:04,300
The council voted on the

2
00:00:04,300 --> 00:00:06,100
proposal this morning.

After

The council voted on the
proposal this morning.

rejoined by hand:
The council voted on the proposal this morning.

The shorter path

When the recording exists, never start from the captions.

Prose written as prose does not need reassembling.

1

Bring the recording

Upload the file, or paste a podcast host, YouTube or Apple Podcasts link. Spotify, Kick VOD, Podbean and Zoom cannot be read and say so at once.

2

Read it on the page

The transcript opens as paragraphs with word-level timestamps; speaker labels show here while you read.

3

Export the text

TXT for plain words, DOCX as a real Word file, MD for notes, JSON for structure. SRT and VTT come from the same job.

Quoting accurately

Checking a quote is far easier in running text, and timestamps in the text let you jump back to the moment.

Still need captions too

Take both. The caption generator covers captions, and SRT vs VTT explains which file to hand over.

Pricing

Text costs the same as captions: nothing extra.

Formats are not priced separately, so a job giving you TXT gives DOCX, SRT, VTT, JSON and MD too. a pack from $4 opens the full transcript in full at any length; after that, packs of 1,000 minutes at $8, 2,000 at $12 or 5,000 at $24, with no expiry.

TXT and MD DOCX for Word JSON with structure From $8 per 1,000 min

FAQ

SRT to TXT, answered.

What exactly gets removed to turn an SRT into text?

Three kinds of line. The counter above each cue, the timecode line containing the arrow, and the blank line separating one cue from the next. What remains is the spoken words, still broken at the points a caption had to break.

Does the free subtitle tool on this site produce a TXT file?

No, and it is worth saying plainly. That tool reads and writes caption formats only, so its output menu offers SubRip and WebVTT and nothing else. For plain text, either export TXT from a transcript or strip the file by hand in an editor.

How do I strip the timecodes by hand?

Open the .srt in any editor with regular expression search. Delete lines matching a pattern for the arrow timecode, delete lines that contain only digits, then collapse the repeated blank lines. It takes about a minute and needs no installation.

Why does the result read so badly?

Because captions are broken for reading speed on screen, not for grammar. A sentence is often split across two or three cues, so stripping the timings leaves short fragments that need rejoining before the text reads as prose.

Is there a way to get clean prose without any of this?

Yes, when you still have the audio or video. Transcribe the recording and export TXT, DOCX or MD directly. The text is written as continuous paragraphs rather than reassembled from caption fragments, so nothing needs repairing afterwards.

Will the text file say who was speaking?

No. A TXT download carries the words alone, and an SRT never held speaker information to begin with. Speaker labels are shown in the transcript view on the page, and of the download formats only JSON records them.

What are people doing with the plain text version?

Usually reading, searching or reusing. A word count for a script, a quote checked against what was actually said, a summary drafted from a talk, or the text pasted into a document where timecodes would only be noise.

How long a recording can I transcribe for free?

A pack from $4 opens the full transcript. Files can run to 10 hours or 5 GB, and anything past that first free transcript draws on credits which do not expire.

Skip the stripping. Export the text.

Drop the recording or paste a link, and take TXT, DOCX and MD from the same transcript.