Audio and podcasting decision
VoiceJournal.io
A single competent developer can build a useful self-hosted VoiceJournal replacement in a few weeks using existing open-source speech and diarization projects; you’ll trade commercial polish, hosted scale, and support for lower costs and control.
Visit website↗Built by Max Hamal 🇺🇦, who ships 8 products in this index
Not priced
No pricing page we fetched carried a figure, so there is nothing to compare against. The build side is still real.
$100one-off60 h to build
$50/mo4 h/mo upkeep
No published price to break even against.
No open-source build does this yet
Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.
What a replacement has to do
- Record or upload audio → transcribe speech → segment & diarize → summarize/convert into note items → store and retrieve notes
What it still won’t have
- Hosted, polished UI and onboarding
- Managed scaling, monitoring, and uptime guarantees
- Any proprietary model tweaks, analytics dashboard, and paid integrations
- Commercial support and product polish
What remains hard
- Product polish and ongoing maintenance
First-year cost
No published price
VoiceJournal.io does not publish a price we could read, so there is nothing to compare against. What building costs is below.
Money you would actually spend
Time you would spend
—
What you would spend
What we assumed
The verdict above measures whether you could build it. This one is only about money.
Runnable build prompt
Build a self-hosted speech-to-notes web app using: Next.js frontend, FastAPI backend (Python), Postgres for storage, and whisper.cpp or faster-whisper for transcription; include a web audio recorder and upload API that accepts MP3/WAV/M4A, run offline transcription to produce text with timestamps, run speaker diarization or whisperX for speaker segments, generate short summaries/section headings (use a local LLM or rule-based heuristics), store transcripts and summaries in Postgres, and provide a UI with transcript playback synced to audio and search. Out of scope: multi-tenant billing, mobile native apps, enterprise SSO. Include error handling, retries for transcription jobs, unit/integration tests for API endpoints, and deployment scripts (Docker Compose).
How we checked
How the score was reached
- Partly verdict base52
- Evidence score52
The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.
How scoring works →Cited sources · 1
Every page the run actually retrieved.
- official productVoiceJournal — Turn speech into perfect notes
Integrity checks
What held up, and what did not.



