Drop a recording (files over 500 MB continue on the upload page). Every line of the transcript is tied to its moment in the audio, with word-level timing underneath, so a quote can be found, checked and cited in seconds. Your first transcript is free, at any length · instant preview with no signup · packs from $8 for 1,000 minutes, credits never expire.
How it works
Nothing to configure. If you only need the words, the audio to text page covers the basics.
A lecture, interview, podcast or voice memo. Up to 10 hours or 5 GB per file; the box here takes files up to 500 MB and larger ones continue on the upload page.
Each segment gets a start and end time and each word carries its own timing. Speaker labels are added as well.
Click a timestamp to hear the line, search for a phrase, then download SRT or VTT for captions, JSON for code, or TXT and DOCX for notes.
How the timing works
Timestamps turn a transcript into an index of the recording.
Each segment shows its start time. Click it and the player starts there, so checking a quote takes one click rather than dragging through an hour of audio.
Every word also carries its own timing. That matters when a sentence start is not precise enough, such as cutting a short clip or lining text up with sound in your own tool.
Search the words in a transcript and go straight to the match. Your transcripts stay in a searchable library, so last month's recording is still one search away.
When you download TXT or DOCX you choose whether each line starts with a bracketed hh:mm:ss timestamp or reads as plain text.
Which file for which job
All five formats are included on every pack.
The caption file most video editors and players accept: numbered blocks, each with a start time, an end time and the words spoken.
The caption format used by HTML5 video players and many course platforms. Same timing as SRT, laid out slightly differently.
Structured data with start and end times and speaker labels, for developers and researchers who want to measure pauses, build a search index or process timing in a script.
Paste a timestamped line into an article draft, meeting notes or a paper. Anyone with the recording can find the moment, for example (Interview 3, 00:14:32).
Pricing
Your first transcript is free, at any length. After that, $8 buys 1,000 minutes, $12 buys 2,000 minutes and $24 buys 5,000 minutes. Credits never expire: about $0.29 to $0.48 per audio hour.
A 2-hour lecture uses $0.96 of credit on the $8 pack, $0.72 on the $12 pack or about $0.58 on the $24 pack, with word-level timestamps, speaker labels and every export format included.
FAQ
Upload the file to the box on this page. The transcript comes back with a timestamp on every line and word-level timing underneath. Preview with no signup; $0.99 opens your first transcript in full, at any length.
Timing for each individual word, not just each sentence or segment. They let you locate a single word in the audio, which helps with clips, captions and analysis.
Yes. Click the timestamp next to any line and the audio plays from that moment, which is the quickest way to check a quote before you use it.
Yes. For TXT and DOCX downloads you choose whether lines start with a timestamp or not. SRT, VTT and JSON always carry timing, because that is what those formats are for.
Both are caption files with start and end times. SRT is the older, widely supported format for video editors; VTT is the web format used by HTML5 players. Whipscribe exports both.
No. $8 for 1,000 minutes, $12 for 2,000 minutes or $24 for 5,000 minutes, about $0.29 to $0.48 per audio hour, with timestamps included. Credits never expire.
Up to 10 hours or 5 GB per file. The box on this page takes files up to 500 MB; larger files continue on the upload page.
Yes. Timestamps make show notes and chapter lists easier to write. See transcribe a podcast for pasting an episode link instead of a file.
Instant preview, no signup, no card. A free account keeps your transcripts in one searchable library.