Learning and careers decision

Spark Playground

A technically competent developer can build a usable, smaller replacement (browser editor + sandboxed Spark runner + question store), but reproducing the hosted product's managed instant Spark execution, polished UX, and paid premium experience is operationally heavier and better suited to continuing to pay.

Visit website
You pay

Not priced

No pricing page we fetched carried a figure, so there is nothing to compare against. The build side is still real.

You’d pay instead

$100one-off80 h to build

$80/mo10 h/mo upkeep

No published price to break even against.

No open-source build does this yet

Nothing published replaces this one, so a replacement starts from an empty file. Here is what it would have to cover.

What a replacement has to do

  • User opens a browser code editor, writes PySpark code against included datasets, sends code to a sandboxed Spark runner, receives execution output and pass/fail feedback against exercise assertions, repeats with new questions.

What it still won’t have

  • Managed, instant browser Spark clusters and scaled execution
  • Polished UX, onboarding flows, and integrated tutorials
  • Any proprietary or curated premium question set and community features
  • Hosted payments, analytics and usage tracking

What remains hard

  • Product polish and ongoing maintenance
Read the build prompt

First-year cost

No published price

Spark Playground does not publish a price we could read, so there is nothing to compare against. What building costs is below.

Money you would actually spend

Keep paying
—

Subscription price × seats × 12

Build it
—

AI build —APIs + hosting —

Time you would spend

—

—

What you would spend

What we assumed

The verdict above measures whether you could build it. This one is only about money.

Runnable build prompt

Not run yet
Build a minimal Spark Playground clone using Next.js (TypeScript) for the frontend, Monaco editor for code editing, a Postgres database for questions and user data, MinIO for sample dataset storage, and a sandboxed Spark runner implemented as Docker containers orchestrated by Docker Compose (or lightweight Kubernetes). Core features in scope: browser editor that submits PySpark scripts; backend run API that launches a containerized Spark job with timeout and resource limits; question CRUD and bundled sample datasets; simple assertion/grading that compares job output to expected results and returns logs; email-based auth and a single paid-gate flag (no full billing integration required). Out of scope: multi-tenant autoscaling clusters, advanced telemetry, marketplace/community features, and high-availability deployment. Include error handling for job timeouts, container failures, and input validation; add unit tests for the grader and integration tests for the run API; provide Docker Compose and a README with local dev and production deploy steps.
How we checked2 sources · 3/3 runs agreed · evidence score 57

How the score was reached

  • Partly verdict base52
  • 2 cited sources+1
  • 3/3 assessment runs agreed+4
  • Evidence score57

The base comes from the verdict. Everything under it is a check that either happened or did not, and each one is a fact frozen in this record rather than a judgement made at render time - so the same evidence always produces the same number.

How scoring works →

Cited sources · 2

Every page the run actually retrieved.

Integrity checks

What held up, and what did not.

✓ 3 independent runs, one answer✓ Citations limited to fetched pages! 1 moat recorded