How well does it handle Chinese (Mandarin)?
Accuracy is strong on clear audio; give names and technical terms a quick scan.
Simplified-script output by default, with sentence segmentation despite no spaces in the script.
Tone-dependent homophones resolve from context. Mandarin is the model's training target; for Cantonese use our dedicated Cantonese page.
What you get
Speaker labels
Every voice in the recording is separated and labelled — works in Chinese (Mandarin) the same as in English, because diarization listens to voices, not words.
Word-level timestamps
Each word is timed to the audio, which is what makes accurate Chinese (Mandarin) SRT and VTT subtitle files possible.
Search & quote
The transcript is searchable in Chinese (Mandarin) — jump to the exact second something was said and copy it with its timestamp.
Every export
TXT, SRT, VTT, DOCX and JSON on every pack — no format paywall.
Who transcribes Chinese (Mandarin) audio here
Meetings, lectures, podcast archives, drama clipping.
Questions people ask
Do I need to tell it the audio is Chinese (Mandarin)?
No — the language is detected from the first seconds of audio automatically. If a recording mixes Chinese (Mandarin) and English, each part is transcribed in the language actually spoken.
Can I get English text from Chinese (Mandarin) audio?
Yes. Alongside transcription in Chinese (Mandarin), the engine can translate the speech to English text in the same pass — useful for subtitling Chinese (Mandarin) content for an international audience.
Does it work on Chinese (Mandarin) video, not just audio?
Yes — upload MP4/MOV/WebM or paste a link; the audio track is extracted automatically and transcribed the same way.
How much does it cost?
The preview is instant with no signup. Credits are one-time purchases that never expire: $4 for 300 minutes, $12 for 2,000 minutes (about 36¢ per hour of audio).
Is this live captioning?
No — Whipscribe transcribes recordings, not live speech. Upload a file that already exists and the Chinese (Mandarin) transcript is typically ready in about a minute per half hour of audio.
Related: Audio to text · Subtitle generator · Interview transcription · Bulk transcription