WhipScribe turns podcasts, interviews, lectures, voice memos, meetings and dictation into text for $0.29 to $0.48 per audio hour, with packs from $4 for 500 minutes. An hour of audio takes about 2 minutes to process, and most files finish in under a minute.
What this is
Most audio transcription services sell one of two things. Human services send your file to a typist, charge by the minute of audio and return it in hours or days. Software tools give you machine transcription inside a monthly subscription with a cap on minutes or file imports. WhipScribe is the second kind, sold the first way: machine transcription you pay for by the minute, in packs that never expire, with no subscription required.
You upload a recording or paste a link, and the transcript is ready while you make a coffee. You read it in the browser with timestamps, search it, and export it. If you transcribe ten hours this month and nothing next month, you pay for ten hours once.
We want to be plain about what it is not. Nobody listens to your file. There is no human proofreading pass, no accuracy guarantee and no certified transcript. For many recordings that does not matter; for some it does, and further down we say when you should pay a human service instead.
What people send us
The same service handles each of these. What changes is what you do with the transcript afterwards.
Paste the episode's RSS or show link, or upload the MP3 you published. Use the transcript for show notes, an accessible episode page or quotes for social posts. The brief and chapters with timestamps come with it.
Research interviews, journalist interviews, oral histories, user research. Timestamps let you jump back to the exact moment a quote was said and check it against the audio before you publish or code it.
A full lecture becomes searchable notes; a term of lectures becomes a study archive. Long recordings stay in one piece: one file can run up to 10 hours, so there is no splitting and stitching.
iPhone Voice Memos (M4A), Android recorder files and voice notes from messaging apps. Good for ideas captured on a walk, drafts dictated in the car, or a message you need in writing.
Olympus and other dictation recorders save DS2 or DSS files. Upload them as they are, with no conversion. If the DS2 file is password-protected, you enter the password when you upload.
Upload the recording from Zoom, Teams, Google Meet or a phone call after it ends. We transcribe recordings, not meetings while they are still going on; for a bot that joins your calls, see the comparison below.
Formats and sources
If a phone, recorder, editor or streaming platform produced it, the chances are good that it uploads as it is. One file can be up to 10 hours long or 5 GB in size.
Every format we support has its own page in the format directory, with any quirks noted. A large uncompressed WAV can pass 5 GB over several hours; save it as MP3 or M4A first and it shrinks to a fraction of the size without hurting the transcript.
What it costs
One-time packs, no seat fees, and credits never expire. The summary of any recording is free to read; a pack opens the full transcript.
| Pack | Minutes | Hours of audio | Per audio hour |
|---|---|---|---|
| $4 pack | 500 | 8 h 20 min | $0.48 |
| $8 pack | 1,000 | 16 h 40 min | $0.48 |
| $12 pack | 2,000 | 33 h 20 min | $0.36 |
| $24 pack | 5,000 | 83 h 20 min | $0.29 |
| Workspace | $19 a month, fair use of 700 hours | — | |
Full details, yearly options and the API are on the pricing page. For a cost-per-hour comparison across services, see cheap transcription services.
Turnaround and what you get
About 2 minutes of processing per hour of audio. Most files finish in under a minute; a full 10-hour recording takes about 20 minutes. There is no queue for a typist and no rush fee.
TXT for notes, DOCX for reports and editing, SRT and VTT for captions, JSON with timestamps for your own tools. All included with every transcript.
Each recording gets a brief with the thesis and takeaways, chapters with timestamps, the names, numbers and dates that were said, and a chat over the transcript.
Languages are detected automatically across 99+ of them, so a Spanish interview or a French lecture needs no setting. Transcripts are kept in a searchable library. Recordings are transcribed on our own servers and are never used for training.
Sample transcripts
These are public recordings, not customer files. Open one and judge the transcript for yourself.
A public-domain audiobook reading: long-form narration with timestamps throughout.
A conversational tech podcast with several hosts, cross-talk and technical vocabulary.
French audio, with the language detected automatically. No setting chosen before upload.
Summaries, chapters and quotes from public podcasts and talks. Two hour-long examples are linked below.
Example write-ups: UK government chief data officer interview (53 min) · TPU origins conversation (57 min)
Honest comparison
Four products that all produce transcripts, built for different jobs. Prices are from each company's own pricing page.
| WhipScribe | Rev | Otter | Descript | |
|---|---|---|---|---|
| What it is | Machine transcription of files and links | Machine transcription plus a human transcription service | Meeting assistant that records and transcribes calls | Audio and video editor driven by the transcript |
| How you pay | Packs from $4 for 500 min; credits never expire | AI plans from $25.49 per seat a month billed yearly ($29.99 monthly); human $1.99 per minute | Pro $16.99 a month, or $8.33 a month billed yearly | Paid plans from $16 a month |
| Included time | Whatever you buy | Essentials: 5,000 AI minutes per user a month | Pro: 1,200 minutes per user a month | Hobbyist: 10 media hours a month |
| Free tier | Summary of any recording free to read | 45 AI minutes a month | 300 minutes a month, 30 minutes per conversation, 3 file imports in total | 1 media hour a month |
| Uploading files | Up to 10 h or 5 GB per file, no import cap | Supported | Pro: 10 imports a month, 90 minutes per conversation | Upload into the editor |
| Human option | No | Yes, 99%+ accuracy guaranteed, 12 hours or less | No | No |
| Languages | 99+, auto-detected | English on Free; English and Spanish on Essentials; 37+ on Pro | Multi-language support | 25 for transcription |
| Where it wins | Lowest cost per hour for occasional or bulk audio; long files; dictation formats | When you need a human-checked or legal-formatted transcript | Live meetings: it joins calls and tells speakers apart as they talk | Editing a podcast or video by editing its text |
If you already pay for Otter or Descript for their main job, the transcription inside them may be all you need. Otter is built for meetings as they happen, which we do not do. Descript is an editor, and its transcript is the way you cut audio, not a document you hand off. Rev offers something no machine service can: a person who listens to every word and stands behind the accuracy.
Where WhipScribe wins is on audio files you already have. There are no monthly minutes to use up and no per-file import limit, so a backlog of interviews or a term of lectures costs the same whether you transcribe it this week or spread it over a year. More detail: vs Rev · vs Otter · vs Descript · best transcription services.
Be honest with yourself
Machine transcription is good enough for most podcasts, lectures, voice memos and interviews you will read and quote yourself. Pay a human service when:
Expect to pay around $1.99 per audio minute for human transcription (Rev's published rate), against under half a dollar per hour here. A practical middle path many people use: machine-transcribe everything, then send only the recordings that matter most to a human service. Our guide to human transcription services covers the options.
How it works
Drop a file in the box at the top of this page or paste a podcast, YouTube or Google Drive link. No software to install.
The free summary shows what the recording covers. Buy a pack from $4 to open the full transcript.
Search the text, jump to any timestamp, and export TXT, DOCX, SRT, VTT or JSON.
FAQ
Packs are $4 for 500 minutes, $8 for 1,000, $12 for 2,000 and $24 for 5,000. That is $0.29 to $0.48 per hour of audio, and credits never expire. The summary of any recording is free to read; a pack opens the full transcript.
About 2 minutes of processing per hour of audio, and most files finish in under a minute. A 10-hour recording takes about 20 minutes.
MP3, M4A, WAV, FLAC, OGG, AAC, OPUS, AIFF, WMA, AMR, GSM, Olympus DS2 and DSS dictation files and most other formats a recorder or phone produces, plus video files. One file can be up to 10 hours or 5 GB. You can also paste a podcast, RSS, YouTube or Google Drive link.
No. WhipScribe is machine (AI) transcription. It is fast and cheap, but nobody listens to your file. For court records, verbatim legal work or very poor audio where every word must be checked, a human service is the better choice.
No. WhipScribe does not label speakers. You get the full text with timestamps. If you need each line attributed to a named person, use a tool that offers speaker identification or a human service.
A transcript with timestamps that you can read and search in the browser, exported as TXT, DOCX, SRT, VTT or JSON. Each recording also gets a brief, chapters with timestamps and a chat over the transcript.
Recordings are transcribed on our own servers and are never used for training. You can delete a recording and its transcript at any time.
99+ languages, detected automatically. You do not need to pick the language before uploading.
Related services
About 2 minutes per audio hour. Packs from $4 for 500 minutes, and credits never expire.
Sources for the comparison, checked 3 Oct 2026 on each company's own page: Rev, rev.com/pricing; Otter, otter.ai/pricing; Descript, descript.com/pricing. Prices are in US dollars, change often and may differ by region; check the source before you buy. Rev, Otter and Descript are trademarks of their owners and are named here only for comparison. WhipScribe prices are from our pricing page.