Send in the video or the link and download SRT and VTT files with cue timings taken from every word. French speech comes back as French cue text; the French transcript page covers the words on their own. Free for your first recording of up to 3 hours, then $8 covers 1,000 minutes.
The short version
Browser only — no plugin, no local install.
Video or audio, uploaded or fetched from a link. Where a recording is longer than the box accepts, the upload page takes it as far as 10 hours or 5 GB.
Every word is timed, so scrubbing to a moment is instant. Proper nouns are worth a pass — Beauvais, Guadeloupe, Saint-Étienne.
SRT for editors and platform uploads, VTT for the web. The same job also gives TXT, DOCX, JSON and MD if the text is needed elsewhere.
Captioning French
Elision is the first one. French glues short words to the next with an apostrophe — l’équipe, d’accord, qu’il, jusqu’à — and that pair is one unit for a reader. A cue line ending just after the apostrophe strands a lone l’ at the edge of the frame, so break before the group.
Then there is spacing punctuation. French typography puts a space before the semicolon, colon, question mark and exclamation mark, and inside guillemets. That space is supposed to be a narrow non-breaking one. Some renderers treat it as an ordinary space and wrap there, dropping a lone ? onto a second line.
The downloaded SRT and VTT hold only cue text and timing — no name tags in front of the lines. Who spoke is visible in the transcript on screen, and carried as data in the JSON export you can pull from the same video job.
Sidecar or burned in
One export, four common destinations.
In Studio, add the file under French rather than letting the automatic track stand. An uploaded cue file is indexed, which puts the spoken words into search.
TikTok and Reels render text into the picture, so bring the SRT into your editor. French lines run long; plan for two.
Resolve and Premiere import SRT as a caption track with the timings intact.
VTT is what a browser expects on a track element. If the video already lives online, start from its link.
Pricing
a pack from $4 opens the full transcript, caption files included. Beyond that the packs are one-time purchases: $8 for 1,000 minutes, $12 for 2,000 minutes and $24 for 5,000 minutes, and unused credits stay on the account.
| French recording | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 8-minute explainer | $0.06 | $0.05 | $0.04 |
| 35-minute documentary cut | $0.28 | $0.21 | $0.17 |
| 90-minute panel from a Paris event | $0.72 | $0.54 | $0.43 |
Figures show how much of a single pack each recording consumes. No format costs extra. The pricing page has the rest.
FAQ
French. The cue text is in the language of the recording, so French speech becomes French captions. Turning them into another language is not something this does.
They stay together as one unit. The cue text keeps the apostrophe attached to the word that follows it, which is how a reader expects to see it on screen.
A renderer has treated the French space before the mark as an ordinary breaking space. Preview a cue that ends in a question, and shorten it if the player wraps awkwardly.
Yes, the « and » characters go into the cue text as they were transcribed, along with accented capitals such as À and É.
It does not. VTT and SRT downloads are cue text plus timings. Speaker separation lives in the transcript view, and in the JSON export as structured data.
Up to 10 hours or 5 GB in one file. Without credits you can submit recordings up to 3 hours, and the first of those is free.
Preview instantly without an account. Your first recording of up to 3 hours costs nothing.