How well does it handle Cantonese?
Supported end-to-end; treat the output as a fast first draft and plan a review pass — it still beats typing from scratch.
Cantonese is supported as its own language (yue) — distinct from Mandarin — with traditional-script output.
Colloquial Cantonese particles (啦, 喎, 咩) appear in the output. Accuracy trails Mandarin; broadcast speech does noticeably better than street conversation, so plan a review pass for anything critical.
What you get
Speaker labels
Every voice in the recording is separated and labelled — works in Cantonese the same as in English, because diarization listens to voices, not words.
Word-level timestamps
Each word is timed to the audio, which is what makes accurate Cantonese SRT and VTT subtitle files possible.
Search & quote
The transcript is searchable in Cantonese — jump to the exact second something was said and copy it with its timestamp.
Every export
TXT, SRT, VTT, DOCX and JSON on every pack — no format paywall.
Who transcribes Cantonese audio here
Hong Kong sermons, radio archives, family-history tape.
Questions people ask
Do I need to tell it the audio is Cantonese?
No — the language is detected from the first seconds of audio automatically. If a recording mixes Cantonese and English, each part is transcribed in the language actually spoken.
Can I get English text from Cantonese audio?
Yes. Alongside transcription in Cantonese, the engine can translate the speech to English text in the same pass — useful for subtitling Cantonese content for an international audience.
Does it work on Cantonese video, not just audio?
Yes — upload MP4/MOV/WebM or paste a link; the audio track is extracted automatically and transcribed the same way.
How much does it cost?
The preview is instant with no signup. Credits are one-time purchases that never expire: $4 for 300 minutes, $12 for 2,000 minutes (about 36¢ per hour of audio).
Is this live captioning?
No — Whipscribe transcribes recordings, not live speech. Upload a file that already exists and the Cantonese transcript is typically ready in about a minute per half hour of audio.
Related: Audio to text · Subtitle generator · Interview transcription · Bulk transcription