Spoken Wikipedia Corpora
Long-form Wikipedia audiobook recordings in English / German / Dutch — ~1000h.
Drop your audio. Transcript in seconds. First transcript free, then $2 a file or $8 = 1,000 min
Best for long-form ASR with article-level context; alignment research. Pricing: free.
What it is
Spoken Wikipedia Corpora (SWC) collects volunteer-read Wikipedia articles in English (~395h), German (~386h), Dutch (~165h). All aligned at sentence + word level. License: CC BY-SA.
Watch out for: CC BY-SA (Wikipedia text + audio) · variable speaker count per article. Cite: Köhn et al., LREC 2016.
Install / use
Go to Spoken Wikipedia Corpora nats.gitlab.ioWhat people actually do with Spoken Wikipedia Corpora-style transcription
The tool is the means. These are the jobs — each one priced at published rates, each one wired up on its own page.
Features
| Speaker diarization | No |
| Word-level timestamps | Yes |
| Streaming / real-time | No |
| Languages supported | 3 |
| HIPAA eligible | No |
Spoken Wikipedia Corpora vs Whipscribe
| Feature | Spoken Wikipedia Corpora | Whipscribe |
|---|---|---|
| Category | Open source | Transcription APIs |
| Pricing | free | $8–$24 one-time packs (credits never expire) · $2 single unlock · free instant preview |
| Speaker diarization | No | Yes |
| Word timestamps | Yes | Yes |
| Streaming | No | No |
| Languages | 3 | 99 |
| Platforms | Web | Web, API, MCP |
Where this category is heading
From the vendor changelogs we track weekly — what changed in August 2026, and what it means if you are choosing now.
AssemblyAI moved summarisation onto an LLM this month; every vendor is racing to return action items, quotes and topics with the text rather than as an add-on.
Whipscribe today Every Whipscribe job already returns an insights payload — summary, key quotes, topics and speakers — from the same job id, at no extra charge.
Deepgram shipped self-hosted container images in August — the market is moving toward audio that stays inside a boundary the customer controls, because teams with customer calls or unreleased material are refusing shared model endpoints.
Whipscribe today Whipscribe runs on our own GPUs in a private, secured cloud. Audio is never forwarded to OpenAI or any third-party model.
The fastest-growing way to use a transcription API is not a form — it is Claude, Cursor or a workflow runner calling it mid-task through MCP.
Whipscribe today Whipscribe ships an MCP server: transcribe, search and summarise from an assistant without wiring anything.
AssemblyAI's 1.0 SDK unified async, realtime and sync; Deepgram's CLI went to 0.3. The unit of work is becoming the folder or the bucket, not the file.
Whipscribe today Submit with an Idempotency-Key and a batch_id, poll by job, retry safely. The S3 connector runs a whole prefix in one grant.
Deepgram added Afrikaans, Georgian and Armenian and improved a dozen more this month. Coverage is widening while quality still clusters around English and the large European languages.
Whipscribe today 99+ languages auto-detected. Ask for a language explicitly when you know it — auto-detect on a short or noisy clip is the most common cause of a wrong-language transcript.
Source: Deepgram and AssemblyAI changelogs, scanned 2026-08-24.
Alternatives to Spoken Wikipedia Corpora
Frequently asked about Spoken Wikipedia Corpora
Is Spoken Wikipedia Corpora free?
Yes. Spoken Wikipedia Corpora is free and open source.
Does Spoken Wikipedia Corpora work on Mac, Windows and Linux?
Spoken Wikipedia Corpora runs in the browser, so it works on any desktop OS with a modern browser, and on phones and tablets.
How do I install Spoken Wikipedia Corpora?
Download the installer for your platform from https://nats.gitlab.io/swc/. No terminal needed.
How many languages does Spoken Wikipedia Corpora support?
Spoken Wikipedia Corpora lists 3 languages.
Does Spoken Wikipedia Corpora identify different speakers?
No. Spoken Wikipedia Corpora does not label speakers; a conversation comes back as one continuous text. If you need speaker labels, that is a separate tool or a different service.
Does Spoken Wikipedia Corpora give word-level timestamps?
Yes — Spoken Wikipedia Corpora produces timing per word, which is what subtitle cues and clip boundaries need.
Can Spoken Wikipedia Corpora transcribe live audio?
No — Spoken Wikipedia Corpora works on finished files, not a live stream.
Is Spoken Wikipedia Corpora HIPAA compliant?
Spoken Wikipedia Corpora does not claim HIPAA compliance. For PHI, run it on infrastructure you control or choose a service that will sign a BAA.
What are the limitations of Spoken Wikipedia Corpora?
CC BY-SA (Wikipedia text + audio) · variable speaker count per article. Cite: Köhn et al., LREC 2016.
Who is Spoken Wikipedia Corpora best for?
Long-form ASR with article-level context; alignment research.
Does Spoken Wikipedia Corpora run offline?
Spoken Wikipedia Corpora runs on your own machine — it is open source, so nothing leaves the computer unless you configure it to.
Is there a Spoken Wikipedia Corpora desktop app?
No — Spoken Wikipedia Corpora is browser-only. There is nothing to install and nothing runs on your machine.
Who makes Spoken Wikipedia Corpora?
Spoken Wikipedia Corpora is made by University of Bielefeld.
What kind of tool is Spoken Wikipedia Corpora?
In this directory Spoken Wikipedia Corpora is filed under dataset as a open-source tool.
What are the alternatives to Spoken Wikipedia Corpora?
There is a side-by-side page at /tools/ksc-spotter-bench-alternatives comparing Spoken Wikipedia Corpora with the closest tools in the same category on price, platform and features.
Whipscribe is a managed faster-whisper + whisperX service. If you want transcripts without running infrastructure, paste a URL or drop a file in the form below — you'll have a transcript in seconds.
Explore
All transcription tools · Audio technology hub · Transcribe any platform · Audio & video formats · How-to guides · Glossary · Playbooks · Apps · Broadcast & radio · Podcast transcripts · Use cases · Blog · Transcription API · Integrations · Automations · For your industry