AI assistants and search decision

Speaky

A technical user can build a usable live speech→slides replacement covering Freestyle and basic Guided modes, but reproducing the vendor's low-latency, highly polished experience (and ongoing product polish) is costly and operationally heavier than a small self-hosted project.

Visit website

Built by Alberto Ravasini ⚡, who ships 5 products in this index

You pay

$5/mo

$60/yr

Read off the official pricing page.

You’d pay instead

$100one-off80 h to build

$150/mo6 h/mo upkeep

On cash alone, building overtakes the subscription at 32 seats.

No open-source build does this yet

Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.

What a replacement has to do

  • Speak → realtime speech-to-text → intent/layout extraction → render slide visuals/overlays → save transcript and slide history

What it still won’t have

  • Sub-second, Cerebras-backed inference speed and the claim of slides appearing "before you finish your sentence"
  • Polished commercial UX, weekly product updates, and the vendor's support/guarantees
  • Any proprietary optimizations for instant overlays and trigger-word responsiveness

What remains hard

  • Product polish and ongoing maintenance
Read the build prompt

First-year cost

Keep paying

Paying is—cheaper in year one.

On cash alone, building overtakes the subscription at 32 seats.

Paid seatsseats

Money you would actually spend

Keep paying
—

Subscription price × seats × 12

Build it
—

AI build —APIs + hosting —

Time you would spend

—

—

What you would spend

What we assumed

The verdict above measures whether you could build it. This one is only about money.

Runnable build prompt

Not run yet
Build a self-hosted minimal live voice-to-slides service using Next.js + React for the frontend, a Node.js/Express WebSocket server, Postgres for storage, and Redis for transient state. Integrate a cloud ASR (e.g., Whisper/any streaming ASR API) for real-time transcript and OpenAI/GPT-style model calls (or local LLM) for layout intent extraction. Core features in scope: live speech streaming with transcript, mapping utterances to 8 layout types, server-side render of slide frames (HTML/SVG) and instant numeric overlays, upload and triggerable mockups, manual slide editor and save/playback of presentations, basic email auth and per-user storage. Out of scope: training custom models, matching Cerebras-level sub-second latency optimizations, multi-tenant billing, analytics dashboard, and mobile app. Include error handling for ASR/model failures, tests for the WebSocket endpoints and core intent extraction, and Docker Compose for local deployment.
How we checked3 sources · 3/3 runs agreed · evidence score 62

How the score was reached

  • Partly verdict base52
  • 3 cited sources+3
  • Price verified on pricing page+3
  • 3/3 assessment runs agreed+4
  • Evidence score62

The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.

How scoring works →

Cited sources · 3

Every page the run actually retrieved.

Integrity checks

What held up, and what did not.

✓ Price read off the page✓ 3 independent runs, one answer✓ Citations limited to fetched pages! 1 moat recorded