Image and video decision
Captioner
A capable developer can reproduce a usable subtitle transcription+editor and burn-in pipeline (using Whisper + FFmpeg and the cited prior-art), but replicating Captioner's polished editor, prioritized rendering, multi-language translation features, and hosted convenience would take substantially more product and ops work.
Visit website↗Built by Simon Liang, who ships 3 products in this index
$10/mo
$120/yr
Read off the official pricing page.
$100one-off78 h to build
$40/mo6 h/mo upkeep
On cash alone, building overtakes the subscription at 5 seats.
No open-source build does this yet
Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.
What a replacement has to do
- Upload video → transcribe audio with an STT model → align timestamps and generate word-level timestamps → edit subtitles in a web editor → export SRT/VTT or burn-in video.
What it still won’t have
- Polish and UX refinements from a product-focused editor (priority rendering, smooth in-browser editing)
- Any proprietary fine-tuning or model optimizations Captioner claims (they say “we use the highest quality AI model” and fine-tuned Whisper for Cantonese)
- Integrated translation workflow and built-in translated language options
- Convenient credit-based billing, account management, and priority rendering offered by the hosted product
What remains hard
- Product polish and ongoing maintenance
First-year cost
Keep paying
Paying is—cheaper in year one.
On cash alone, building overtakes the subscription at 5 seats.
Money you would actually spend
Time you would spend
—
What you would spend
What we assumed
The verdict above measures whether you could build it. This one is only about money.
Runnable build prompt
Build a minimal self-hosted captioning web app using React for the frontend, Node.js + Express for the API, PostgreSQL for user/asset metadata, S3-compatible storage for uploads, and a worker queue (BullMQ) to run transcription and rendering tasks. Core features in scope: secure upload of MP4/MOV, run Whisper (local or via hosted STT) to produce transcripts with word-level timestamps, timestamp alignment and chunking into SRT/VTT, a web subtitle editor to edit text and timing, SRT/VTT export, and an FFmpeg-based background job to burn-in burned-in MP4s. Out of scope: multi-tenant billing/crediting system, priority rendering tiers, advanced translation UI. Require error handling for failed transcriptions and renders, retry logic in the worker queue, API tests for upload/transcription/export flows, and basic end-to-end tests for the editor export functionality.
How we checked
How the score was reached
- Partly verdict base52
- 3 cited sources+3
- Price verified on pricing page+3
- 3/3 assessment runs agreed+4
- Evidence score62
The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.
How scoring works →Cited sources · 3
Every page the run actually retrieved.
- official productCaptioner Home
- official pricingCaptioner Pricing
- official docsCaptioner Supported Languages
Integrity checks
What held up, and what did not.


