Transcript to SRT: text has no timing, so start from the recording.

An SRT is words plus the moment each line is spoken. A TXT or Word transcript only has the words. Upload the recording itself: you get a timed SRT in minutes, and your existing transcript becomes the reference for corrections.

Timing from the audioYour wording via correctionsSRT and VTTFree preview

How it works

From a plain transcript to a timed SRT in three steps.

1

Upload the recording

The audio or video your transcript was made from. Timing can only be measured from the sound.

2

Bring your wording across

Compare the new transcript with yours and correct lines where they differ. Edits keep the cue's timing.

3

Download the SRT

The file has your wording where you changed it and measured timing throughout.

Why a converter cannot do this

TXT to SRT is not a format change.

Converting SRT to VTT is a format change: the same information, written differently. Converting TXT to SRT is not, because the SRT needs something the text file does not contain: when each line is said. A tool that offers to convert text to SRT without the audio either asks you to type every timecode, or spreads the lines evenly across a duration you give it, which is wrong by seconds within the first minute.

So which situation are you in?

Keeping a transcript you trust

When the existing text is the approved version.

Sometimes the transcript is a document someone has signed off: a court-reporter's text, a published interview, a script. You want that wording, timed. The practical method:

  1. Transcribe the recording here to get cues with timing.
  2. Open your approved text beside the transcript and correct each line that differs. Names, numbers and specialist terms are where the differences will be.
  3. Download the SRT. Timing is measured; wording is yours.

For a scripted video where the speaker follows the script closely, the differences are few and this takes minutes. YouTube also offers its own alignment: upload a plain text file in Studio with the Without timing option and it will time the lines for supported languages.

Sample output

What the SRT file looks like.

1
00:00:34,680 --> 00:00:46,500
fourscore and seven years ago our fathers brought forth upon this continent a new nation conceived in in liberty, and dedicated to the proposition

2
00:00:46,500 --> 00:00:50,300
that all men are created equal.

3
00:00:50,300 --> 00:00:57,500
Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived

4
00:00:57,500 --> 00:01:01,620
and so dedicated, can long endure.

5
00:01:01,620 --> 00:01:07,920
We are met on a great battlefield of that war we have come to dedicate a portion of that field

6
00:01:07,920 --> 00:01:14,240
as a final resting place for those who here gave their lives that their nation might live

7
00:01:15,360 --> 00:01:18,800
it is altogether fitting and proper that we should do this

Real output, unedited: a public-domain LibriVox reading of the Gettysburg Address (2 min 41 s) transcribed here in October 2026. The repeated word and the missing capitals are what came back; you would fix those on the transcript before downloading.

Before you rely on it

What the subtitle file is, and what it is not.

One cue per spoken phrase

Each cue is a phrase from the transcript with its start and end time. Most run a few seconds. A passage spoken without a pause can come out as one long cue of up to about half a minute, so split those in your editor if your platform wants two short lines at a time.

Timing comes from the audio

Cue times are taken from the recording itself, so the file lines up with the video or audio you uploaded. If you trim the start of the video afterwards, shift the cues by the same amount in your editor.

Accuracy we measured

On public test recordings, about 7 words in 100 differed from the reference for clean English read speech, between 4 and 9 in 100 for Spanish, Portuguese, Russian, French and German, and about 21 in 100 for Hindi. Noisy rooms, music and crosstalk do worse. Method and every number: the benchmark.

Fix a line before you download

Open the transcript, correct a misheard name or number on its line, and the SRT and VTT you download carry the correction with the timing unchanged.

A plain file, not a styled video

You get the subtitle file. Fonts, colours, position and captions drawn into the picture are done in your video editor or player, which reads this file.

Limits

Up to 10 hours or 5 GB per recording. The box on this page takes files up to 500 MB; larger ones continue on the upload page. 99+ languages, detected automatically.

Pricing

What timing a transcript costs

The preview is free. When the full transcript needs a pack, one-time packs are $4 for 500 minutes, $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes. Credits never expire, which comes to about $0.29 to $0.48 per hour of audio.

Recording$4 or $8 pack$12 pack$24 pack
5-minute scripted video$0.04$0.03$0.02
45-minute interview$0.36$0.27$0.22
3-hour hearing$1.44$1.08$0.86

Figures are the share of a pack each recording uses. A button opens secure checkout for that pack. Every export format is included; full details on the pricing page.

FAQ

Transcript to SRT, answered.

Can I convert a TXT or Word transcript to SRT?

Not by conversion alone, because a text transcript has no timing. Upload the recording it came from to get a timed SRT, then correct the wording to match your transcript.

I have the transcript and the audio. What is the fastest way to get an SRT?

Upload the audio here. You get a timed SRT in minutes, and you use your transcript only to correct the lines that differ.

Can it align my exact text to the audio automatically?

No. It transcribes the audio and gives you timed cues that you can edit to your wording. For automatic alignment of a script, YouTube Studio's upload with the Without timing option does this for supported languages.

My transcript has timestamps already. Do I need the audio?

Not necessarily. A subtitle editor such as Subtitle Edit can import timestamped text and turn it into SRT; you set where each line ends.

How do I go from SRT back to plain text?

Use the SRT to TXT converter, which removes numbering and timecodes in the browser, or download TXT from the transcript.

Subtitles and captions

Every page in this set uses the same upload box and the same export. All subtitle pages

Closest to this page

Make subtitles

From a link or a file

WebVTT

For a platform or an editor

Convert a subtitle file

Formats and rules

By language

Upload the recording. Get the timing.

Free preview. Timed SRT in minutes; packs from $4.