Upload the video or paste a link and download Japanese SRT and VTT files whose cues come from word-level timings. Japanese audio returns Japanese cue text. Only need the text? The Japanese transcript page is the one. packs from $4, credits never expire, and 1,000 minutes cost $8 afterwards.
The process
In the browser, with no cue typed by hand.
Upload a file or paste a link. Longer masters continue on the upload page, which handles up to 10 hours or 5 GB.
Every word is timed, so navigating is instant. Names and specialist terms are the place to spend correction time, since kanji choice matters.
SRT for editing suites and platform uploads, VTT for browsers, with TXT, DOCX, JSON and MD from the same job.
Captioning Japanese
Japanese is written without spaces between words, which removes the natural break points a Latin-script caption relies on. A renderer that wraps wherever the line runs out will happily split a word, or separate a particle from the phrase it marks. Japanese captioning practice breaks at phrase boundaries instead — after a particle such as は, が, を or に, or between clauses — so each line is a unit a reader can take in at a glance.
Character density changes the line budget completely. Broadcast practice in Japan has long worked to roughly thirteen full-width characters a line, two lines at a time, because each character carries far more information than a Latin letter and takes about twice the width. Porting an English cue sheet's forty-two-character lines into Japanese produces lines nobody can read in the time available.
Nothing in the downloaded SRT or VTT labels the speaker — both are cue text and timecodes. The transcript view separates voices on screen, and the JSON export from the same run keeps that as data.
Delivery
Platform track, editor import, or burned into the picture.
Attach the SRT as the Japanese track in Studio, which replaces the automatic captions and makes the words searchable.
Text is burned in on these platforms. Import the SRT into your editor and set a line length suited to full-width characters.
Both import SRT onto a caption track with timings intact; choose a font with full Japanese coverage before the master render.
VTT loads through an HTML5 track element. For published video, start from the link instead of the master file.
Pricing
A summary is free and readable in full at any length, and the caption files come with it. After that: $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes, one-time, credits that do not expire.
| Japanese video | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 10-minute product video | $0.08 | $0.06 | $0.05 |
| 45-minute 対談 | $0.36 | $0.27 | $0.22 |
| 105-minute seminar | $0.84 | $0.63 | $0.50 |
Each cell is the fraction of one pack the recording uses. No format costs more than another; see the pricing page.
FAQ
Japanese. Cue text is written in the language that was spoken, so Japanese audio gives Japanese captions and nothing is converted into another language.
At a phrase boundary — after a particle or between clauses — rather than wherever the line happens to run out.
Far fewer than an English line holds. Japanese broadcast practice works around thirteen full-width characters per line over two lines, because each character is wider and denser.
Yes, full-width punctuation is written as transcribed, and lines are not started with a closing bracket or a small kana.
SRT and VTT are horizontal caption formats, so the cues are laid out horizontally. Vertical text would be a styling decision in a separate tool.
No. That belongs to the transcript view and the JSON export; SRT and VTT hold cue text and timings only.
Preview instantly with no account. The opening transcript is free for anything up to 3 hours.