Drop a field recording, a Zoom MP4, or paste a URL. Multi-speaker transcript with timestamps every 30 seconds, one-click quote extraction, DOCX export. Built for journalists, oral historians, qualitative researchers. Packs from $4, credits never expire. · instant preview with no signup · packs from $4 for 500 minutes, credits never expire.
How it works
Built for the journalism + research workflow — speaker fidelity, citation timestamps, and the formats your editor or thesis advisor expects.
Field-recorder WAV, phone-recorded M4A, Zoom MP4, or a hosted URL. Up to 10 hours or 5 GB per file (up to 3 hours before you buy credits); files over 500 MB continue on the upload page.
Every voice is labelled automatically. The transcript view marks each turn as Speaker 1 and Speaker 2, so you can see which voice is the reporter and which is the source as you read.
One click for a quote list — strongest statements, attributed, timestamped. One click for a DOCX hand-off ready for your editor or thesis chapter.
What you can do with the transcript
Every interview transcript opens with an AI sidebar tuned for long-form Q&A.
Pull-quote candidates surfaced automatically — attributed to a speaker, timestamped, ready for a deck card or a story headline.
Word-level timestamps on every line. Hover any sentence for the exact second; click to jump back to the audio.
Ask “every time the source mentions funding” or “list all named individuals.” The AI answers with quoted timestamps.
DOCX for an editor, TXT for a quick draft, JSON for a research database. SRT/VTT if you're cutting an explainer video around the audio.
Pricing
Packs from $4, credits never expire. After that, a 1-hour profile interview costs 48¢ on the $8 pack and a 3-hour roundtable about $1.44. Prorated to the second; no monthly subscription; credits never expire.
FAQ
Drop the audio or video file above, or paste a URL. Multi-speaker transcript with timestamps returns in about a minute per 30 minutes of audio. Preview with no signup; a pack from $4 opens the full transcript.
Yes. Every voice is separated automatically and shown as Speaker 1 and Speaker 2 in the transcript view on the page. Three-plus voices handled the same way.
Yes. Whipscribe stamps every 30 seconds by default for skimmability, with word-level timestamps on every line for citation precision.
Yes. One click after the transcript loads produces a quote list with speaker attribution and timestamps.
Yes. Open-source Whisper on private infrastructure. Audio is never sent to OpenAI, Google, or any third-party transcription service, and we never train models on your audio. TLS 1.3 in transit, AES-256 at rest.
MP3, M4A, WAV, FLAC, MP4, MOV, MKV. Field-recorder WAV from Zoom H1n/H6 works directly.
Yes. DOCX with speaker labels and timestamps included on every pack. TXT, SRT, VTT, JSON also included.
Instant preview, no signup, no card. A free account keeps your transcripts in one library.