AI assistants and search decision
THE BLUE HOUSE
A competent developer can assemble a usable local dictation tool using whisper.cpp and related OSS within a few weeks, but reproducing Weesper's cross‑app polish, platform optimizations, and frictionless subscription/update UX is substantial—so building a narrow self-hosted replacement is realistic, cloning the full commercial product less so.
Visit website↗$5/mo
$60/yr
Read off the official pricing page.
$100one-off90 h to build
$0/mo3 h/mo upkeep
On cash alone, building overtakes the subscription at 2 seats.
No open-source build does this yet
Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.
What a replacement has to do
- User holds global hotkey → capture microphone audio → run local ASR model to transcribe → post-process (punctuation, filler removal, language detection/translation) → insert resulting text at the active cursor.
What it still won’t have
- Production polish (HUD responsiveness, <1s latency on Apple Silicon) and performance tuning
- Prebuilt Apple‑Silicon optimizations and compact model packaging
- Seamless cross‑app insertion edge-cases and broad app compatibility testing
- Built-in subscription flow and Stripe integration packaged inside the app
- Polished contextual prompts UI and curated custom-dictionary UX
What remains hard
- Execution quality
Metal optimized (Apple Silicon) — Maximum performance on Mac M1/M2/M3/M4
- Execution quality
Discreet HUD overlay — Audio visualization during dictation
First-year cost
Keep paying
Paying is—cheaper in year one.
On cash alone, building overtakes the subscription at 2 seats.
Money you would actually spend
Time you would spend
—
What you would spend
What we assumed
The verdict above measures whether you could build it. This one is only about money.
Runnable build prompt
Build a cross-platform desktop dictation app (Electron frontend + Rust backend). Use whisper.cpp for local ASR inference and ship three local model sizes (small/medium/large) selectable in settings. Implement: global hotkey audio capture, on-device transcription with <2s latency on Intel and <1s on Apple Silicon (optimize via Metal/FFI), post-processing pipeline (punctuation, filler removal, language detection, optional translation), universal text insertion using macOS and Windows accessibility/input APIs, a minimal HUD showing live audio level and status, local storage for custom prompts/dictionaries, and an update checker. Out of scope: hosted transcription, multi-user cloud sync, enterprise licensing. Include robust error handling, unit tests for core pipeline components, and CI that builds macOS and Windows installers.
How we checked
How the score was reached
- Partly verdict base52
- 2 cited sources+1
- Price verified on pricing page+3
- 3/3 assessment runs agreed+4
- Evidence score60
The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.
How scoring works →Cited sources · 2
Every page the run actually retrieved.
- official productWeesper Neon Flow — product page
- official pricingWeesper Neon Flow — pricing
Integrity checks
What held up, and what did not.

