Chinese (Mandarin) documentary interviews and field footage, transcribed properly
Chinese (Mandarin) (中文(普通话)): Simplified-script output by default, with sentence segmentation despite no spaces in the script. Tone-dependent homophones resolve from context. Mandarin is the model's training target; for Cantonese use our dedicated Cantonese page.
Documentary footages in particular: Timecoded text is the paper edit: find the moment in text, cut in the timeline. Searchable rushes, pull-quote selects and subtitle files for screeners.
How it works
Record well
Feed the raw interview files; multi-hour material is expected.
Upload it
Drop the file into Whipscribe or paste a link. 30 minutes free daily, no signup to try.
Read & export
Speaker-labeled, timestamped Chinese (Mandarin) text in minutes. Export TXT, SRT, VTT, or DOCX.
Typical uses: Meetings, lectures, podcast archives, drama clipping.
Frequently asked
Can Whisper transcribe Chinese (Mandarin) documentary interviews and field footage?
Yes. Simplified-script output by default, with sentence segmentation despite no spaces in the script. Tone-dependent homophones resolve from context. Mandarin is the model's training target; for Cantonese use our dedicated Cantonese page.
How are speakers handled in a documentary footage?
Timecoded text is the paper edit: find the moment in text, cut in the timeline.
How much does it cost?
About 3.3¢ per audio minute (~$2 an hour), 30 minutes free every day, no signup to try.
Is my recording private?
Yes — private infrastructure, never sent to a third-party AI service, no training on uploads. Always get participants' consent before recording.