Who said what · upload & go

Transcription with speaker labels: see who said what.

Drop an interview, meeting or panel recording (files over 500 MB continue on the upload page). The transcript comes back split into turns, each marked with a numbered speaker label and a timestamp you can click to hear that moment. Packs from $4, credits never expire. · instant preview with no signup · packs from $4 for 500 minutes, credits never expire.

Speaker labels on every turn Word-level timestamps TXT/DOCX/SRT/VTT/JSON Up to 10 hours or 5 GB per file

How it works

From a recording to a labelled transcript in three steps.

Nothing to install and no setting to switch on: speaker labels come with every transcript.

1

Upload the recording

A two-person interview, a team call or a panel. Up to 10 hours or 5 GB per file; the box here takes files up to 500 MB and larger ones continue on the upload page.

2

Voices are separated

The transcript is split each time the voice changes. Every turn carries a numbered label for the speaker, plus word-level timestamps.

3

Check and export

Click a timestamp to hear any line, search for a word, then download TXT, DOCX, SRT, VTT or JSON.

What you get

A labelled transcript reads like a script.

A plain transcript is one long block of words. With speaker labels, a conversation keeps its shape.

Turns, not paragraphs

Each time a different person starts talking, a new block begins with that speaker's label and start time. A 45-minute interview reads as a back-and-forth you can scan in minutes.

Numbered, not named

Voices can be told apart by sound, but the transcript cannot know who anyone is. Labels are numbered, so keep a short note of who spoke first to match labels to people before you quote.

Hide labels for a clean read

A Show speakers switch in the transcript view turns the labels off when you want a plain reading copy, and back on when you need to attribute a quote.

Labels and timing as data

The JSON export keeps the speaker next to the start and end time of every segment. Use it to total each person's talk time or feed your own tools.

Honest limits

When speaker labels struggle.

Separating voices works best on clear recordings. These are the cases to check by clicking the timestamp and listening.

Crosstalk

When two people talk over each other, the words usually come through, but they can land under whichever voice was louder. Check heated moments and quick interruptions.

Similar voices

Two speakers with a similar pitch and pace, especially sharing one microphone, can be merged under one label or swapped part-way through a recording.

Very short turns

One-word replies such as yes, right or mm-hm give very little voice to go on, so they are the most likely to be attributed to the wrong speaker.

The room and the microphone

A microphone per person, less echo and fewer people talking at once all help. One phone in the middle of a table of eight is the hardest case.

Pricing

Speaker labels are included at the base price.

Packs from $4, credits never expire. After that, $8 buys 1,000 minutes, $12 buys 2,000 minutes and $24 buys 5,000 minutes. Credits never expire, which works out to about $0.29 to $0.48 per audio hour.

A 1-hour interview uses $0.48 of credit on the $8 pack. A 3-hour panel uses $1.44 on the $8 pack, $1.08 on the $12 pack or about $0.86 on the $24 pack. Full details are on the pricing page.

$8 for 1,000 min $12 for 2,000 min $24 for 5,000 min Credits never expire

FAQ

Speaker labels, answered.

What is transcription with speaker labels?

It is a transcript that marks where each speaker's turn begins, so you can see who said what. Whipscribe adds numbered speaker labels and word-level timestamps to every transcript at no extra cost.

How do I get a transcript with speaker labels?

Upload the recording to the box on this page. Labels are added automatically, with nothing to switch on. Preview with no signup; a pack from $4 opens the full transcript.

Can speech to text tell speakers apart?

Yes, by how each voice sounds. That means it can number speakers but not name them. Clear recordings with distinct voices separate best; crosstalk and similar voices are harder.

Why are two people under the same label?

Similar voices, a shared microphone, people talking over each other or very short replies can merge speakers. Click the timestamp of the line in question and listen before you attribute a quote.

Does it work for interviews and meetings?

Yes. Two-person interviews are the easiest case; meetings and panels work too, with more checking where people interrupt. For interview tips see transcribe an interview.

Can I correct the transcript?

On your own transcripts you can edit the wording of a line or delete a line in the transcript view, then save. Downloads then use your edited version.

What does it cost?

$8 for 1,000 minutes, $12 for 2,000 minutes or $24 for 5,000 minutes, about $0.29 to $0.48 per audio hour, with speaker labels included. Packs from $4, credits never expire, and credits never expire.

How long can the recording be?

Up to 10 hours or 5 GB per file. The box on this page takes files up to 500 MB; larger files continue on the upload page.

Drop a recording. See who said what.

Instant preview, no signup, no card. A free account keeps your transcripts in one searchable library.