WebM is what the browser itself records: screen captures from web tools, camera recordings made on a page, audio grabbed by a web app. Drop one in and download cues timed from word-level timestamps. Above 500 MB the upload page handles it, as far as 10 hours or 5 GB. Packs from $4 any length · the preview appears with no signup · packs open at $8 per 1,000 minutes, no expiry.
How it works
A recording made on the web gets captioned on the web.
Whatever the recorder handed you. Up to 500 MB in the box above, with longer captures going through the upload page.
The speech comes back with a timestamp on every word, usually a few minutes of waiting per hour captured, and no account is needed to see it.
VTT is the caption format browser players load, which suits a file that was recorded in a browser. SRT, TXT, DOCX, JSON and Markdown come too.
WebM specifics
WebM turns up in two shapes. Most commonly it holds picture and sound together, from a screen or camera capture; less often it holds audio only, which is what a page recording from a microphone tends to write. Either uploads the same way, and the difference only decides what you do with the cues afterwards — burn them into the picture, or lay them over whatever the audio ends up inside.
Screen recordings compress efficiently in WebM because so much of the frame holds still; an hour of a shared screen often lands in the hundreds of megabytes rather than the gigabytes an equivalent camera recording would take. A long capture of full-motion video is another matter, and for anything approaching 5 GB it is quicker to upload the sound alone using the free extract audio from video tool, which preserves the timeline exactly.
This one, for a .webm of either shape. For picture in another container, video to SRT; for a sound-only file, audio to SRT; for a public link, YouTube to SRT; for anything else, the caption generator. Container details are on the WebM format page.
Where the file goes
VTT and WebM are both web-native, so a recording made in a browser can be published with captions in the same toolkit.
A screen capture explaining a feature gains readable words, which is what makes it useful to someone watching without sound.
Timed text lets a colleague read what was said at a given moment instead of replaying the capture.
SRT imports into Premiere Pro and DaVinci Resolve, so a browser recording can be cut and captioned like any other footage.
Pricing
Packs from $4 and readable in full at any length. After that, 1,000 minutes cost $8, 2,000 cost $12, 5,000 cost $24, with credits that do not lapse.
| Recording | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 10-minute demo | $0.08 | $0.06 | $0.05 |
| 50-minute walkthrough | $0.40 | $0.30 | $0.24 |
| 2-hour workshop | $0.96 | $0.72 | $0.58 |
| 4-hour capture | $1.92 | $1.44 | $1.15 |
Figures are the share of a pack each recording uses, by runtime. See pricing.
FAQ
Drop the file into the box on this page, preview it with no signup, open the full transcript with a free account, then download SRT or VTT.
Not at all. An audio-only WebM is ordinary input, and cues are built from speech rather than from frames.
For a browser player, yes, since VTT is what it loads. SRT is the better pick for editing software and video platforms, and both come from the same job.
No. Speaker labels appear in the transcript you read on screen. SRT, VTT, TXT and DOCX downloads hold words and timings only; the JSON export is the one with speaker information.
Upload its audio instead, using the free extract audio from video tool. The timeline is unchanged, so the cues still match the recording frame for frame.
An hour costs $0.48 from the $8 pack, or $0.29 if you buy the $24 one. A capture can run to 10 hours or 5 GB.
Instant preview, no signup. A free account opens the first capture in full, to 3 hours.