Speech to text · in the browser

Speech to text from any recording, in your browser.

Drop a recording of people talking (files over 500 MB continue on the upload page). You get the words back as text, split into timed segments with speaker labels, ready to read, search or export. Packs from $4, credits never expire. · instant preview with no signup · packs from $4 for 500 minutes, credits never expire.

Speaker labels included Word-level timestamps TXT/DOCX/SRT/VTT/JSON Up to 10 hours or 5 GB per file

How it works

Recorded speech in, text out.

This page is for speech you have already recorded: a meeting, lecture, interview, sermon or voice note.

1

Upload the recording

An audio or video file up to 10 hours or 5 GB. The box here takes files up to 500 MB; larger ones continue on the upload page.

2

Speech becomes text

The words come back as timed segments, with a numbered label for each speaker and word-level timestamps.

3

Read, search, export

Click any line to hear it, search for a word, then download TXT, DOCX, SRT, VTT or JSON. Transcripts stay in your library.

A practical guide

Better text starts with the recording.

Most transcript problems come from how the speech was captured, not from the file format. Four habits make the biggest difference.

Record close to the speaker

Distance matters more than microphone price. A phone a hand's width from the speaker usually beats an expensive microphone on the far side of the room.

Cut the background

Music, air conditioning and café noise all compete with the voice. If you can choose the room, choose the quiet one with soft furnishings.

One voice at a time

Overlapping speech is harder to turn into clean text and harder to attribute. In meetings, a little turn-taking goes a long way.

Keep the original file

Upload the original recording rather than a copy squeezed through a messaging app, which can strip detail from the voice.

What to expect

Where a quick human check still pays off.

Speech to text saves the typing, not the reading. Skim these before you publish or quote.

Names and jargon

People's names, product names and field-specific terms are the words most often spelled wrong. A quick search for each one finds every instance.

Crosstalk and short replies

When people talk over each other, or answer in a single word, the words and speaker labels there deserve a listen.

Speech to text or audio to text?

They describe the same conversion. This page is the practical guide for spoken recordings; audio to text covers converting audio files in general.

Specific jobs

Recording an interview or a podcast? Transcribe an interview and transcribe a podcast have tips for those.

Pricing

Pay per minute of speech, not per month.

Packs from $4, credits never expire. After that, $8 buys 1,000 minutes, $12 buys 2,000 minutes and $24 buys 5,000 minutes. Credits never expire: about $0.29 to $0.48 per audio hour.

A 30-minute voice note uses $0.24 of credit on the $8 pack. An 8-hour recording, such as a full conference day, uses $3.84 on the $8 pack, $2.88 on the $12 pack or about $2.30 on the $24 pack.

$8 for 1,000 min $12 for 2,000 min $24 for 5,000 min Credits never expire

FAQ

Speech to text, answered.

What is speech to text?

Software that turns spoken words in a recording into written text. Whipscribe returns that text in timed segments, with speaker labels and word-level timestamps.

How do I convert speech to text online?

Upload the recording to the box on this page and the transcript appears in your browser, with nothing to install. Preview with no signup; a pack from $4 opens the full transcript.

Is speech to text free?

Packs from $4, credits never expire. and anyone can preview a transcript with no signup. After that, packs start at $4 for 500 minutes and credits never expire.

Can it handle more than one speaker?

Yes. Each turn is marked with a numbered speaker label. Labels are most reliable when voices are distinct and people take turns.

What files can I use?

Audio or video recordings up to 10 hours or 5 GB per file. The box on this page takes files up to 500 MB; larger files continue on the upload page.

How do I get better results?

Record close to the speaker, in a quiet room, with one person talking at a time, and upload the original file rather than a compressed copy.

What can I export?

TXT, DOCX, SRT, VTT and JSON, all included on every pack. SRT and VTT are caption files; JSON keeps the timing data for code.

Does it keep my transcripts?

Yes. With a free account your transcripts are kept in a searchable library, so you can find a recording by its file name or its words later.

Drop a recording. Read it as text.

Instant preview, no signup, no card. A free account keeps your transcripts in one searchable library.