OpenSuperWhisper

by Starmel · maintained fork by My-Monkey

A free, MIT-licensed Mac dictation app. Hold a shortcut, speak, let go, and the text lands in whatever window you were already typing in — with three of its four engines running entirely on your own machine.

macOS 14 Sonoma or later  ·  free

Have a recording to get through today?

Drop your audio. Transcript in seconds. Packs from $4, readable in full, any length

TL;DR

A free, MIT-licensed macOS dictation app. You hold a shortcut or a spare mouse button, speak, release, and the words are inserted into whichever app has focus. Four engines are available and three of them run fully on your Mac — nothing is uploaded, there is no account, and there is no telemetry. It also takes dropped audio files and has a small command-line entry point.

Best for people who dictate all day on a Mac and want that loop to be local, private and free. It needs macOS 14 Sonoma or later, and it is a dictation tool first: a recording you need as subtitles, a document, or a transcript you can share is a different job.

Category
Open source
License
MIT
Stars
★ 2.9k
Latest release
v0.12.4 · 2026-09-12
Pricing
free
Platforms
macOS 14+

What it is

OpenSuperWhisper is a Swift app that turns your Mac into a dictation machine. Bind a key combination, a single modifier such as Right Option, or a thumb button on your mouse; hold it, talk, let go, and the text is typed into Slack, your editor, a browser field — whatever was in front of you. An indicator shows the recording, and with one engine you can watch the words appear as you speak.

Recognition runs locally through whisper.cpp — the open-source speech model family — alongside two other on-device engines and an optional remote one. Models download from inside the app and load lazily, so browsing the settings never triggers a surprise download.

Best for: Mac users who dictate constantly and want it private, local and free.
Watch out for: macOS only, two active repositories under one name, and no sense of who is speaking.

Install / use

brew install --cask my-monkeys/tap/opensuperwhisper
Use the full tap path. There are two live projects with this name. The original, Starmel/OpenSuperWhisper, has ~2.9k stars and is still being worked on, but its last tagged release is 0.1.0 from 3 March 2026 — and that is the build the bare opensuperwhisper Homebrew cask still installs. The community fork at my-monkeys/OpenSuperWhisper shipped v0.12.4 on 12 September 2026 with notarized disk images for both Apple Silicon and Intel. Typing the bare name is the most common way people end up on a nine-month-old build.

Four engines, three of them local

You pick per model, and rules can switch the model automatically depending on which app you are dictating into — a fast one for chat, an accurate one for email.

whisper.cppOn deviceAccuracy, around 99 languages, translation into English
Parakeet (FluidAudio)On deviceSpeed and a live preview, 25 European languages
SenseVoice (sherpa-onnx)On device · Apple Silicon onlyChinese, Cantonese, English, Japanese, Korean
RemoteA server you chooseAny endpoint you point it at, with a local fallback if it is unreachable

The extras around the edges are the reason people stay: a custom dictionary so your proper nouns come out spelled right, per-app formatting rules, optional filler-word removal, an opt-in cleanup pass that can run a small on-device model for punctuation and casing, a history you can cap or switch off entirely, and a command line — opensuperwhisper transcribe note.m4a prints text to stdout, or JSON with --json.

Where it is genuinely better than a hosted service

Three things, and they are not small. It is free, with no per-hour cost and no account to create. Nothing leaves your Mac on the local engines, so audio you would never upload — a client call, a medical note, an unreleased recording — never has to travel. And the dictation loop is fast in a way a website cannot be: the text appears where your cursor already is, with no tab switch, no upload, no copy and paste. If you dictate for hours a day on a Mac you own, this is the right shape of tool and we would rather say so than pretend otherwise.

Where it gets awkward

It is macOS only, and only macOS 14 or later — no Windows build, no Linux build, nothing on a phone, and nothing for a machine you cannot install software on. Models download before the first run and the accurate ones are large; the optional cleanup model is roughly another gigabyte. An older Mac takes real time over a long recording, and while it works, it is your machine that is busy.

More to the point, it is shaped around dictation rather than recordings. Dropped files queue up and the command line returns text or JSON, but there is no subtitle editor, no document export, and no transcript page to send to someone. It does not tell you who is speaking, so an interview comes back as one undifferentiated block. And the two-repository situation means a real chance of installing an old build and judging the project by it.

When a hosted transcript is the better answer

Installing anything is the wrong move when the job is one file this afternoon, when you are on a phone or a borrowed laptop, or when you need the result in a particular format within the hour. That is the gap WhipScribe fills: paste a link or upload a file, and the transcript appears in the browser. The summary is free and readable in full, at any length, with no signup to start or to preview it.

Files can run to 10 hours or 5 GB — the widget here takes up to 500 MB, anything larger goes through the upload page. Exports are TXT, DOCX (a real .docx), SRT, VTT, JSON and MD, which matters when the transcript has somewhere to be: an MP3 you need as text, a recording hours long, text with timestamps. Over 100 languages, and the transcript comes back in the language that was spoken. Speaker labels appear in the transcript view; of the downloads, only JSON carries them. Spotify, Kick VODs, Podbean and Zoom meeting links are refused, so bring the file for those.

Pricing is a one-time credit pack, not a subscription: $8 for 1,000 minutes, $12 for 2,000, $24 for 5,000. Credits never expire and there are no seats. Plenty of people use both — OpenSuperWhisper for dictation that should never leave the machine, a hosted run for the recording they would rather not tie up their Mac with.

OpenSuperWhisper vs Whipscribe

FeatureOpenSuperWhisperWhipscribe
CategoryOpen source, localHosted, in the browser
Pricingfree$8–$24 one-time packs, credits never expire · packs from $4
Runs onmacOS 14+ onlyAny browser, phone included
Audio leaves your machineNo on the local enginesYes — the file is uploaded
Install requiredYes, plus a model downloadNo
Live dictation into any appYesNo
Subtitle and document exportsNo — text or JSONTXT, DOCX, SRT, VTT, JSON, MD
Languages~99100+

Links

Alternatives to OpenSuperWhisper

Frequently asked about OpenSuperWhisper

Is OpenSuperWhisper free?

Yes. It is MIT-licensed and free forever, with no Pro tier and no paywall. The maintainers accept optional donations on Ko-fi, and nothing in the app is gated behind them.

Does OpenSuperWhisper run on Windows or Linux?

No. It is a macOS app and needs macOS 14 Sonoma or later. The maintained fork ships builds for both Apple Silicon and Intel Macs, though the SenseVoice engine is Apple Silicon only.

Does it send my audio anywhere?

Three of its four engines run entirely on your Mac and send nothing. The fourth, called Remote, sends audio to a server you configure yourself and labels itself plainly in the interface. There is no account and no telemetry.

Which repository should I install from?

The maintained fork is my-monkeys/OpenSuperWhisper, on v0.12.4 as of 12 September 2026. The Homebrew cask called opensuperwhisper still points at the original project at 0.1.0, so install the fork through its full tap path: brew install --cask my-monkeys/tap/opensuperwhisper.

Can OpenSuperWhisper transcribe a recording rather than live dictation?

Yes. Drag audio files onto the app and they queue up, or run opensuperwhisper transcribe file.m4a from a terminal, which prints the text to stdout, or JSON with the --json flag.

Does it label who is speaking?

No. It returns one stream of text. For an interview or a panel where you need to know who said what, that is not the job this app was built for.

How many languages does it handle?

Around 99 with automatic detection, depending on the engine. Parakeet covers 25 European languages and SenseVoice covers Chinese, Cantonese, English, Japanese and Korean. The app can also translate speech into English as it transcribes.

What does it cost to run compared with a hosted transcript?

Nothing per hour once it is installed. You pay in disk space, download time and the minutes your Mac spends working. A hosted run costs money per hour and no setup: at WhipScribe a pack from $4 opens the full transcript, and credit packs start at $4 for 500 minutes and never expire.

Can I use OpenSuperWhisper in a browser or on a phone?

No. It installs on a Mac. On a phone, a borrowed laptop or a work machine you cannot install software on, you need something that runs in the browser: paste a link or upload a file at WhipScribe and read the transcript on the page.

What tends to go wrong with it?

Models download before the first run and the accurate ones are large, an older Mac grinds through a long recording, the dictation shortcut can collide with another app, and the optional cleanup model is another gigabyte on disk. None of that is unusual for local software. It is simply work you do rather than work someone does for you.

Whipscribe is the hosted side of the same job. Nothing to install and nothing to configure — paste a URL or drop a file below, and you will have a transcript in seconds.

Explore