Learning and careers decision

Elispeak

A technical user can build a useful, narrow replacement (real-time practice + basic feedback) using existing ASR and LLM APIs and prior-art repos, but reproducing Elispeak's polished curriculum, tuned pronunciation models, and product polish would be larger effort.

Visit website
You pay

$9.99/mo

$120/yr

Read off the official pricing page.

You’d pay instead

$100one-off60 h to build

$30/mo6 h/mo upkeep

On cash alone, building overtakes the subscription at 4 seats.

Open-source builds that already do this

Every project below is open source and already does this job today. Fork one, self-host it, or take the parts you need - the build prompt further down assumes an empty file, and this is the shortcut past that. Licences differ; check the one on each card before you ship. All Elispeak alternatives, with the arithmetic →

What a replacement has to do

  • User speaks into the app → audio is transcribed → system analyzes pronunciation/grammar and returns actionable feedback → user repeats corrected utterance and progress is recorded.

What it still won’t have

  • Polished curriculum, topic catalog and UX polish
  • Proprietary pronunciation-analysis models and tuned scoring
  • Mobile app and cross-device polish
  • Marketing, brand and user acquisition
  • Instructor/teacher features and classroom flows

What remains hard

  • Product polish and ongoing maintenance
Read the build prompt

First-year cost

Keep paying

Paying is—cheaper in year one.

On cash alone, building overtakes the subscription at 4 seats.

Paid seatsseats

Money you would actually spend

Keep paying
—

Subscription price × seats × 12

Build it
—

AI build —APIs + hosting —

Time you would spend

—

—

What you would spend

What we assumed

The verdict above measures whether you could build it. This one is only about money.

Runnable build prompt

Not run yet
Build a minimal web-based AI English speaking tutor using Next.js for the frontend, Node.js/Express for the backend, PostgreSQL for storage, and WebRTC/getUserMedia for audio capture. Core features in scope: topic selection, real-time audio capture and upload, ASR integration (call a cloud speech-to-text API), a simple feedback engine that compares transcripts to expected answers and returns grammar/pronunciation hints, repeat/retake flow, minutes/quota accounting, user auth, and a dashboard showing per-session progress. Out of scope: mobile native apps, advanced pronunciation phoneme-level alignment, paid subscription billing integration, and a polished curriculum. Include error handling for network and ASR failures, automated tests for backend endpoints, and Docker deployment scripts for hosting.
How we checked4 sources · 2/3 runs agreed · evidence score 63

How the score was reached

  • Partly verdict base52
  • An open-source build was found+5
  • 4 cited sources+3
  • Price verified on pricing page+3
  • Evidence score63

The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.

How scoring works →

Cited sources · 4

Every page the run actually retrieved.

Integrity checks

What held up, and what did not.

✓ Price read off the page! 2 of 3 runs agreed; the verdict is the majority✓ Citations limited to fetched pages! 1 moat recorded