Upload the recording or paste a link. The interview comes back as Spanish text — the words that were spoken, not an English version — with each turn marked in the transcript view and every word timed, so you can hear a quote again before you print it. Files over 500 MB continue on the upload page. A pack from $4 opens the full transcript · packs from $4 for 500 minutes.
How it works
Two voices or ten, and nothing to install.
Drop the audio or video into the box above, or paste a link to it. The box takes files up to 500 MB; a longer session continues on the upload page, up to 10 hours or 5 GB.
The transcript view marks each turn with a numbered speaker label, so a question and the answer to it sit apart on the page instead of running together in one block.
Click a timestamp to hear the line again. Then take the text away as TXT, DOCX, SRT, VTT, MD or JSON.
Spanish as spoken
Regional variation inside a single conversation is normal, and nothing is set per region beforehand.
A researcher in Mexico City interviewing someone from Madrid is an ordinary recording here. Both voices are transcribed as Spanish, in one pass, with no dialect to choose first.
Argentine and Uruguayan forms — vos tenés, vos sabés, the shifted stress — are written the way the speaker said them rather than tidied into tú.
The distinction peninsular speakers make between ce, zeta and ese, and the seseo of most of Latin America, both come out as ordinary spelling.
Clipped delivery and dropped final consonants from the Caribbean coast are the hardest case. Clear audio matters more there than anywhere else.
What you get back
Fieldwork, reporting and long sit-downs.
Numbered labels separate the turns on screen. TXT, DOCX, SRT and VTT carry the words without them; JSON keeps the speaker on every segment if you need that in your own tooling.
Packs from $4, credits never expire and readable in full, at any length, which covers most one-to-one interviews without buying a pack at all.
Every word is timed to the audio. Before a quote goes into a chapter or a story, click it and listen — misheard names are the usual catch.
Overlap is where any transcript struggles. The stretch is still written out, but read it against the recording.
Pricing
A pack from $4 opens the full transcript. After that, one-time packs: $8 for 1,000 minutes, $12 for 2,000 minutes and $24 for 5,000 minutes. Credits never expire.
Twelve interviews of fifty minutes come to 600 minutes — comfortably inside the $8 pack, with 400 minutes left over for the follow-ups.
FAQ
Yes. The transcript is what was said, written in Spanish. It is not turned into English — Spanish audio produces Spanish text, and that is the file you download.
In the transcript view, yes: each turn carries a numbered speaker label, so questions and answers sit apart as you read. The labels are numbers rather than names.
No. TXT, DOCX, SRT and VTT contain the words without speaker labels. Only the JSON export keeps the speaker on every segment.
That is fine. There is no regional setting to choose, and a recording with a Mexican voice and a peninsular voice is handled as one Spanish interview.
Up to 10 hours or 5 GB per file. The box on this page takes files up to 500 MB; anything larger continues on the upload page.
A stretch spoken in English comes back as English text and a stretch spoken in Spanish comes back as Spanish. Nothing is translated in either direction.
DOCX or TXT for reading and coding, MD for a plain-text editor, and JSON when you need the speaker and timing of every segment.
Instant preview, no signup. A pack from $4 opens the full transcript.