Upload the video or paste a link and download Thai SRT and VTT files whose cues follow word-level timings. Thai speech returns Thai cue text. If you want the transcript instead, that page covers it. First recording free at up to 3 hours; packs begin at $8 for 1,000 minutes.
How it goes
Nothing to install, no manual cue timing.
Video or audio, uploaded or from a link. Longer masters continue on the upload page, up to 10 hours or 5 GB.
Each word carries its own timing, so moving around is quick. Names and borrowed technical terms are worth a correction pass.
SRT for editors and platform uploads, VTT for browser players, along with TXT, DOCX, JSON and MD.
Captioning Thai
Thai writes without spaces between words. Where a space does appear, it marks the end of a phrase or a sentence, doing the job a comma or a full stop does elsewhere. That inverts the usual assumption in caption tooling: a renderer that wraps at spaces is not breaking between words at all, and a renderer that wraps anywhere will happily split a word down the middle. Cue lines are built to break at those phrase spaces so a reader gets a complete unit.
The script also stacks. A single Thai syllable can carry a consonant, a vowel sign above or below it, and a tone mark above that vowel, three levels deep. Anything that truncates a line by counting code points can cut between the consonant and its marks, leaving an orphan sign at the start of the next line. Line length in Thai has to be judged by what appears on screen.
The SRT and VTT downloads carry cue text and timings only, with no name in front of any line. Voices are separated in the transcript on screen, and the JSON export from the same job holds that as data.
Delivery
Platform track, burned-in text, or a web player.
Attach the SRT as the Thai track in Studio. It replaces automatic captions and makes the words searchable.
These render text into the frame. Import the SRT into your editor and check that the font stacks vowel and tone marks cleanly.
SRT arrives as a caption track with the timings right; test a stacked syllable on screen before the master render.
VTT loads through an HTML5 track element, where the browser handles Thai shaping. Start from the link for published video.
Pricing
A pack from $4 opens the full transcript, caption exports included. Beyond that, packs are one-time: $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes, and credits do not expire.
| Thai video | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 14-minute tutorial | $0.11 | $0.08 | $0.07 |
| 45-minute interview | $0.36 | $0.27 | $0.22 |
| 110-minute seminar | $0.88 | $0.66 | $0.53 |
Each figure is the part of one pack the recording spends. Formats are never billed apart; see the pricing page.
FAQ
Thai. Cue text follows the spoken language, so Thai audio produces Thai captions and no conversion into another language happens.
They mark the end of a phrase or sentence rather than the end of a word, so they are the right places for a caption line to break.
A tool truncated the line by counting code points and cut between a consonant and its stacked marks. Judge line length by what appears on screen.
Yes. The stacked vowel and tone marks are written into the cue text, and they survive as long as the file stays in UTF-8.
Thai does not use one. The spacing carries that role, which is why the cue breaks follow the phrase spaces.
No. SRT and VTT hold cue text and timings; the transcript view separates the voices and the JSON export records it.
Instant preview without an account. The first recording of up to 3 hours is on us.