Image and video decision
REKALAB
A technical user can build a narrow single-prompt→60s pipeline using open-source models and tooling, but reproducing Rekalab's hosted polish, scale, priority queues, and any proprietary scene/voice quality would be difficult and time-consuming.
Visit website↗$19/mo
$228/yr
Read off the official pricing page.
$100one-off64 h to build
$150/mo6 h/mo upkeep
On cash alone, building overtakes the subscription at 9 seats.
Open-source builds that already do this
Every project below is open source and already does this job today. Fork one, self-host it, or take the parts you need - the build prompt further down assumes an empty file, and this is the shortcut past that. Licences differ; check the one on each card before you ship. All REKALAB alternatives, with the arithmetic →
What a replacement has to do
- 1) Prompt -> script generation via an LLM; 2) Per-scene visual generation (text-to-image) and simple motion (pan/zoom); 3) Text-to-speech voiceover generation; 4) Scene composition into a single MP4 (captions, text animation, transitions) with FFMPEG; 5) Simple web UI + storage for uploads and downloads.
What it still won’t have
- Proprietary scene-generation model and any quality differences in Rekalab's 'Better Scene Model'
- Hosted priority queue, scaling, and polished multi-tenant billing/credits system
- Polished UI/UX and integrated gallery with automatic retention policies
- Commercial voice licenses and any curated voice models Rekalab provides
- Built-in templates, analytics, and marketing integrations
What remains hard
- Execution quality
The world's simplest video creation tool.
- Execution quality
Built fast. Shipped faster.
First-year cost
Keep paying
Paying is—cheaper in year one.
On cash alone, building overtakes the subscription at 9 seats.
Money you would actually spend
Time you would spend
—
What you would spend
What we assumed
The verdict above measures whether you could build it. This one is only about money.
Runnable build prompt
Build a single-tenant web app (Flask or FastAPI backend, React frontend) that converts one prompt into a 60s MP4 using LocalAI + Diffusers for text-to-image and an open-source TTS (Coqui TTS or LocalAI voice model). Core features in scope: accept a single prompt, call an LLM to produce a script and scene breakdown, generate one image per scene, apply simple motion (pan/zoom) and animated captions, synthesize voiceover, compose scenes and audio into a single MP4 via FFMPEG, allow image/video uploads per scene, and provide download. Out of scope: training models, multi-tenant billing, priority queues, social scheduling, analytics dashboard. Include error handling, retries for model calls, unit tests for key backend flows, and a Dockerfile + deployment instructions to run on a single VM with S3-compatible storage.
How we checked
How the score was reached
- Partly verdict base52
- An open-source build was found+5
- 4 cited sources+3
- Price verified on pricing page+3
- Evidence score63
The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.
How scoring works →Cited sources · 4
Every page the run actually retrieved.
- official productRekalab homepage
- official pricingRekalab pricing section
- open sourcecalesthio/OpenMontage
- open sourceHBAI-Ltd/Toonflow-app
Integrity checks
What held up, and what did not.





