Image and video decision

Captioner

A capable developer can reproduce a usable subtitle transcription+editor and burn-in pipeline (using Whisper + FFmpeg and the cited prior-art), but replicating Captioner's polished editor, prioritized rendering, multi-language translation features, and hosted convenience would take substantially more product and ops work.

Visit website

Built by Simon Liang, who ships 3 products in this index

You pay

$10/mo

$120/yr

Read off the official pricing page.

You’d pay instead

$100one-off78 h to build

$40/mo6 h/mo upkeep

On cash alone, building overtakes the subscription at 5 seats.

No open-source build does this yet

Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.

What a replacement has to do

  • Upload video → transcribe audio with an STT model → align timestamps and generate word-level timestamps → edit subtitles in a web editor → export SRT/VTT or burn-in video.

What it still won’t have

  • Polish and UX refinements from a product-focused editor (priority rendering, smooth in-browser editing)
  • Any proprietary fine-tuning or model optimizations Captioner claims (they say “we use the highest quality AI model” and fine-tuned Whisper for Cantonese)
  • Integrated translation workflow and built-in translated language options
  • Convenient credit-based billing, account management, and priority rendering offered by the hosted product

What remains hard

  • Product polish and ongoing maintenance
Read the build prompt

First-year cost

Keep paying

Paying is—cheaper in year one.

On cash alone, building overtakes the subscription at 5 seats.

Paid seatsseats

Money you would actually spend

Keep paying
—

Subscription price × seats × 12

Build it
—

AI build —APIs + hosting —

Time you would spend

—

—

What you would spend

What we assumed

The verdict above measures whether you could build it. This one is only about money.

Runnable build prompt

Not run yet
Build a minimal self-hosted captioning web app using React for the frontend, Node.js + Express for the API, PostgreSQL for user/asset metadata, S3-compatible storage for uploads, and a worker queue (BullMQ) to run transcription and rendering tasks. Core features in scope: secure upload of MP4/MOV, run Whisper (local or via hosted STT) to produce transcripts with word-level timestamps, timestamp alignment and chunking into SRT/VTT, a web subtitle editor to edit text and timing, SRT/VTT export, and an FFmpeg-based background job to burn-in burned-in MP4s. Out of scope: multi-tenant billing/crediting system, priority rendering tiers, advanced translation UI. Require error handling for failed transcriptions and renders, retry logic in the worker queue, API tests for upload/transcription/export flows, and basic end-to-end tests for the editor export functionality.
How we checked3 sources · 3/3 runs agreed · evidence score 62

How the score was reached

  • Partly verdict base52
  • 3 cited sources+3
  • Price verified on pricing page+3
  • 3/3 assessment runs agreed+4
  • Evidence score62

The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.

How scoring works →

Cited sources · 3

Every page the run actually retrieved.

Integrity checks

What held up, and what did not.

✓ Price read off the page✓ 3 independent runs, one answer✓ Citations limited to fetched pages! 1 moat recorded