Drop your MP3 here. Hosted Whisper Large v3 returns a text transcript with word-level timestamps in about a minute per 30 minutes of audio. Exports DOCX, SRT, VTT, TXT, JSON at the base price. Packs from $4, credits never expire · instant preview with no signup · free account adds 60 minutes.
How it works
A hosted GPU backend means no local install, no Python, no CLI.
64 kbps voice memo or 320 kbps studio file — both work. Up to 2 GB per file in the browser uploader. WAV, FLAC, M4A, MP4, MOV also accepted.
30-minute MP3 transcribes in about a minute. Word-level timestamps and automatic speaker diarisation included.
DOCX for a writer handoff, SRT/VTT for captions, TXT for a quick paste, JSON for downstream tooling. All included on every pack.
What you can do with the transcript
Every MP3 transcript opens with an AI sidebar — search, summarize, chat, export.
Click a line to jump to the second in the audio player. Search a name or phrase — every match returns with a timestamp.
One click produces a chapter-marked summary, key quotes, and a short social blurb.
Ask any question of the audio — “what did the speaker conclude?” or “list every action item.” The AI answers with quoted timestamps.
DOCX, SRT, VTT, TXT, JSON — every format included on every pack, no Pro paywall.
Pricing
A 30-minute MP3 is free. A 1-hour file runs about $2. Prorated to the second; no monthly subscription, no credit card to start.
FAQ
Drop the MP3 onto the upload box above. Text transcript with word-level timestamps returns in about a minute per 30 minutes of audio. Preview with no signup; a free account (60 minutes included) unlocks the full text.
No. 64 kbps voice memos through 320 kbps studio files all work. Lower bitrates with heavy background noise may dip in accuracy, but the typical range is handled well.
DOCX, SRT, VTT, TXT, and JSON — all included on every pack. No Pro-tier paywall on any export.
Whisper Large v3 on a hosted GPU backend: 30-minute MP3 in about a minute, 1-hour file in about two minutes.
Yes. Automatic diarisation labels every voice as Speaker 1, Speaker 2, and so on; rename them after the transcript loads.
Up to 2 GB per file in the browser uploader. Larger files can go through the API or paste-a-URL flow.
Yes. We run open-source Whisper on private infrastructure. Audio is never sent to OpenAI, Google, or any third-party transcription service, and we never train models on your audio.
Instant preview, no signup, no card. A free account includes 60 minutes.