Upload the video or paste a link and download Chinese SRT and VTT files whose cues follow word-level timings. Chinese speech comes back as Chinese cue text. If the plain transcript is what you need, that page is here. First recording to 3 hours at no cost, then 1,000 minutes for $8.
How it runs
Nothing installed, nothing timed by hand.
Video or audio in, from a file or a link. Longer recordings continue on the upload page, up to 10 hours or 5 GB.
Timings sit on each word, so scrubbing is quick. Names, places and product terms are where a correction pass pays off.
SRT for editing suites and platform uploads, VTT for browsers, plus TXT, DOCX, JSON and MD.
Captioning Chinese
Chinese writes without spaces between words, so a caption renderer has no built-in signal for where a line may wrap. Left alone it will break between any two characters, including inside a two-character compound, which reads as a typo to a viewer. Cue lines should break at phrase boundaries — after a topic, before a verb phrase, at a comma — so each line holds a unit that can be read whole.
Density sets the line budget. Each character is full-width and carries roughly a word's worth of meaning, so a Chinese caption line holds far fewer characters than an English one while saying about the same thing. Practical practice keeps lines in the mid-teens of characters, single-line where possible, because a viewer reads the whole line as a shape rather than scanning it.
Both caption formats carry cue text and timecodes and nothing else, so no line is prefixed with a speaker name. The transcript view shows the separation between voices, and the JSON export from the same job holds it as structured data.
Delivery
A single export serves platforms, editors and players alike.
Attach the SRT as the Chinese track. An uploaded file is indexed, unlike burned-in text, so the words become searchable.
These render text into the picture. Import the SRT into your editor and keep the line short, since full-width characters fill a line fast.
Both read SRT onto a caption track with timings preserved; pick a font with full character coverage before the final render.
VTT is the caption file a browser reads from a track element. For already-published video, build it from the link.
Pricing
a pack from $4 opens the full transcript, caption files included. Afterwards the packs are single purchases: $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes, and credits never expire.
| Chinese video | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 8-minute 产品介绍 | $0.06 | $0.05 | $0.04 |
| 30-minute 访谈 | $0.24 | $0.18 | $0.14 |
| 180-minute 直播回放 | $1.44 | $1.08 | $0.86 |
Each number is the share of one pack the recording uses. No format is charged separately — see the pricing page.
FAQ
Chinese. Captions are written in the language that was spoken, and they are not turned into any other language.
The characters written are the ones produced from the speech as transcribed. No conversion between character sets is applied on your behalf, so check the file if your audience expects a specific set.
At a phrase boundary — after a topic, before a verb phrase, or at a comma — never in the middle of a two-character compound.
Considerably fewer than an English line carries. Keeping lines in the mid-teens of full-width characters, single-line where possible, reads comfortably.
Yes. The enumeration comma, pause comma and full stop are written as transcribed and each takes a full character slot on the line.
No. They contain cue text and timings only; the transcript view separates voices on screen and the JSON export records it.
The preview is free and needs no account. The first recording of up to 3 hours is free of charge.