Upload the video or paste a link. The subtitles come back in Japanese script — the words as spoken, not an English version — timed from word-level timestamps and ready for an editor. Files over 500 MB continue on the upload page. A pack from $4 opens the full transcript · packs from $4 for 500 minutes.
How it works
The soundtrack is read out of the video for you.
MP4, MOV or WebM into the box above, or paste a public link. Anything over 500 MB continues on the upload page.
Each word is timed against the audio and the cues are built from those times. Your first transcript opens in full at any length.
SRT for editors and platform uploads, VTT for web players, and TXT, DOCX, MD and JSON from the same job.
The character budget
The timing comes out of the job. Fitting the characters is the editorial work that follows.
Japanese subtitle practice commonly allows roughly 16 full-width characters per line and two lines per cue — a far tighter budget than the same scene in English.
Around four characters a second is a widely used ceiling for adult programming. A dense line that is technically on screen long enough can still be unreadable, so trim rather than hold.
Japanese does not separate words with spaces, which means a line break falls wherever you decide a phrase ends. The file gives you timings; the break is an editorial judgement.
、 and 。 and the full-width ? need the file imported and saved as UTF-8, or the cues arrive in the editor as mojibake.
Sidecar or burned in
What you download is the sidecar file; burning in happens in your editor.
SRT or VTT stays separate: the viewer can switch it off, a platform can read the words, and a correction does not mean re-exporting.
For subtitles that cannot be turned off, import the SRT into Premiere Pro or DaVinci Resolve and set the typeface and position there.
SRT and VTT carry words and timings only. If a scene needs the speaker named on screen, add it in the editor — JSON is where the speaker of each segment is kept.
When a line will not fit the budget, cut words rather than reduce the type size. Subtitles that shrink stop being readable on a phone.
Pricing
A pack from $4 opens the full transcript. After that, one-time packs: $8 for 1,000 minutes, $12 for 2,000 minutes and $24 for 5,000 minutes. Credits never expire.
A 10-minute episode is $0.08 on the $8 pack. A hundred of them is 1,000 minutes — one $8 pack for the whole run.
FAQ
Yes, in Japanese script rather than romaji. Nothing is turned into English: Japanese video produces Japanese cues, and that is the file you download.
Common practice is around 16 full-width characters per line and two lines per cue, at roughly four characters a second. The file gives you the timing; fitting the characters is an editing decision.
Every word carries its own timestamp and the cues are built from those, so a line appears when it is spoken rather than on an even split of the running time.
No. Numbered speaker labels appear in the transcript view on screen. SRT and VTT hold words and timings; JSON is the export that keeps the speaker of each segment.
That is almost always an encoding mismatch. Import and save the file as UTF-8 and the kanji, kana and full-width punctuation come through intact.
Public video links are usually fetched and transcribed. Some sources refuse to hand over audio, Spotify, Podbean and Kick VODs among them, and for those upload the file.
Up to 10 hours or 5 GB per file. The box here takes 500 MB, and your first transcript reads in full at any length.
Instant preview, no signup. A pack from $4 opens the full transcript.