Audio and podcasting decision

VoiceJournal.io

A single competent developer can build a useful self-hosted VoiceJournal replacement in a few weeks using existing open-source speech and diarization projects; you’ll trade commercial polish, hosted scale, and support for lower costs and control.

Visit website

Built by Max Hamal 🇺🇦, who ships 8 products in this index

You pay

Not priced

No pricing page we fetched carried a figure, so there is nothing to compare against. The build side is still real.

You’d pay instead

$100one-off60 h to build

$50/mo4 h/mo upkeep

No published price to break even against.

No open-source build does this yet

Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.

What a replacement has to do

  • Record or upload audio → transcribe speech → segment & diarize → summarize/convert into note items → store and retrieve notes

What it still won’t have

  • Hosted, polished UI and onboarding
  • Managed scaling, monitoring, and uptime guarantees
  • Any proprietary model tweaks, analytics dashboard, and paid integrations
  • Commercial support and product polish

What remains hard

  • Product polish and ongoing maintenance
Read the build prompt

First-year cost

No published price

VoiceJournal.io does not publish a price we could read, so there is nothing to compare against. What building costs is below.

Money you would actually spend

Keep paying
—

Subscription price × seats × 12

Build it
—

AI build —APIs + hosting —

Time you would spend

—

—

What you would spend

What we assumed

The verdict above measures whether you could build it. This one is only about money.

Runnable build prompt

Not run yet
Build a self-hosted speech-to-notes web app using: Next.js frontend, FastAPI backend (Python), Postgres for storage, and whisper.cpp or faster-whisper for transcription; include a web audio recorder and upload API that accepts MP3/WAV/M4A, run offline transcription to produce text with timestamps, run speaker diarization or whisperX for speaker segments, generate short summaries/section headings (use a local LLM or rule-based heuristics), store transcripts and summaries in Postgres, and provide a UI with transcript playback synced to audio and search. Out of scope: multi-tenant billing, mobile native apps, enterprise SSO. Include error handling, retries for transcription jobs, unit/integration tests for API endpoints, and deployment scripts (Docker Compose).
How we checked1 sources · 2/3 runs agreed · evidence score 52

How the score was reached

  • Partly verdict base52
  • Evidence score52

The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.

How scoring works →

Cited sources · 1

Every page the run actually retrieved.

Integrity checks

What held up, and what did not.

! 2 of 3 runs agreed; the verdict is the majority✓ Citations limited to fetched pages! 1 moat recorded