Upload the video or paste a link. The captions come back in Spanish — the words as spoken, not an English version — timed from word-level timestamps and ready for an editor or a platform upload. Files over 500 MB continue on the upload page. A pack from $4 opens the full transcript · packs from $4 for 500 minutes.
How it works
The soundtrack is read out of the video for you.
Drop an MP4, MOV or WebM into the box, or paste a public link. Up to 500 MB goes straight in; a feature-length file continues on the upload page.
The lines are timed from the timestamp on each word, so a cue lands when it is spoken.; a pack from $4 opens the full transcript.
SRT for video editors and platform uploads, VTT for players on the web, plus TXT, DOCX, MD and JSON.
Lines people can read
The file gives you the timing. These are the editorial calls that make it readable.
Spanish generally needs more characters than the English of the same dialogue, so cues that felt comfortable on an English cut arrive crowded. Expect to shorten a few lines.
Caption practice usually holds a line to around 42 characters and a cue to two lines. The line breaks are yours to set in a subtitle editor.
Around 17 characters a second is the usual ceiling for adult programming. Where a fast exchange breaks it, split the cue in two.
Tildes, ñ and the opening ¿ and ¡ are written into the file. Import and save as UTF-8 so they survive the trip into your editor.
Sidecar or burned in
What comes out of here is the sidecar. The burn-in happens in your editor.
SRT or VTT stays its own file. A viewer can switch it off, a platform can index the words, and a typo is a two-second fix.
For captions that cannot be turned off, import the SRT into Premiere Pro or DaVinci Resolve and render it into the frame.
SRT and VTT hold words and timings, not speaker labels. If a two-hander needs names on screen, add them in the editor — the JSON export is where the speaker of each segment lives.
Pull the soundtrack out and upload that instead. The timeline is unchanged, so the cues still land correctly against the picture.
Pricing
A pack from $4 opens the full transcript. After that, one-time packs: $8 for 1,000 minutes, $12 for 2,000 minutes and $24 for 5,000 minutes. Credits never expire.
A 20-minute explainer is $0.16 on the $8 pack. A weekly 35-minute show for a year comes to 1,820 minutes, which sits inside the $12 pack of 2,000.
FAQ
Yes. The caption file holds the Spanish that was spoken. Nothing is turned into English: Spanish video produces Spanish cues, and that is what downloads.
Yes, at no extra cost. Every finished transcript exports as SRT, VTT, TXT, DOCX, MD and JSON, so one job serves both a web player and an editor.
Every word carries its own timestamp and the cues are built from those, so a line appears when it is spoken.
No. Speaker labels appear in the transcript view on screen. SRT and VTT carry words and timings only; JSON is the export that keeps the speaker of each segment.
Often, yes — public video links are fetched and transcribed. Some sources refuse to hand over their audio, Spotify, Podbean and Kick VODs among them, and those you upload.
Each stretch is written in the language it was spoken in, so a bilingual video produces a caption file that switches where the speakers do. Nothing is translated.
Up to 10 hours or 5 GB per file. The box here takes up to 500 MB, and your first transcript reads in full at any length.
Instant preview, no signup. A pack from $4 opens the full transcript.