GoldieBench deep dive
Muse Spark 1.2's full benchmark breakdown.
Meta's coding reasoning model — co-trained with its own agent, 1M-token window. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Strong: a beautifully rendered GPU Julia set with smooth coloring, orbit-trap glow, and a polished glassmorphic control panel (mode toggle, iterations, palettes, live zoom/center stats). Weak: nothing major visible — a very complete, on-brief, task-winning build.”
“Strong 3D scene with vivid, silky aurora curtains over snowy dunes, silhouetted pine trees, moon and a frozen lake, plus a polished glassy HUD with Kp badge and control chips. Rich composition and glowing shader ray detail push it past the field's best; only minor nit is the somewhat flat foreground snow expanse.”
“Gorgeous dense rain with bright white heads, proper katakana glyphs, and a cohesive cyberpunk HUD (stream status, carrier signal, control bar) that elevates it well past a generic canvas demo. Strong glow, vignette, and interactivity hooks make it a task winner; the only minor nit is the 3D Three.js layer being mostly hidden behind the dominant rain.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Muse Spark 1.2's lowest scored builds, verdicts unedited:
“The HUD renders cleanly with polished synthwave styling, but the entire 3D scene — road, sun, mountains, car — is a black void, so the actual outrun driving game never appears. A functional shell around a non-rendering canvas fails the core brief.”
“The HUD is polished and thematically on-brief (compass, health/stamina/magicka bars, Cinzel fonts), but the 3D world is bare and unconvincing—flat washed-out terrain, only a couple of low-poly trees and rocks, and a giant floating gray monolith that looks like a broken mesh. No visible enemies, buildings, or sense of an explorable Skyrim world, making it feel unfinished versus the field's best.”
“HUD is genuinely polished — health/fury/speed meters, altitude/rings/combo, score+best, minimap, and control hints all look shippable — but the actual 3D scene is broken: the camera is buried inside a giant cone/horn with no visible dragon body, no neon rings, no canyon, and no sky. The core gameplay view fails to render on-brief.”
“The HUD is polished and on-brief (torch fuel, vitalis, kills, coffin altar), but the rendered scene looks flat and bright rather than torch-lit and dark — lighting is washed-out with pale beige walls, floating spheres, and no atmospheric torch glow, badly undermining the dungeon-crawler mood.”
05 · Category breakdown vs the whole field
Games (23 tasks)
Others (3 tasks)
Pages (3 tasks)
Sims (12 tasks)
Visuals (9 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Muse Spark 1.2 perform on GoldieBench?
Muse Spark 1.2 averages 7.47/10 across 50 scored one-shot build tasks, ranking #15 of 22 ranked frontier models, with 0 task golds, 3 silvers and 4 bronzes. Every score is a real judged build you can open and play on this site.
02What is Muse Spark 1.2 best at?
Its strongest scored build is Fractal at 8.7/10. By category it averages Game 7.0, Other 7.1, Page 8.2, Sim 8.2, Visual 7.5 on the bench.
03How much does Muse Spark 1.2 cost?
Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field.
Source ledger
- 01Official vendor site: developer.meta.com/ai/models/muse-sparkdeveloper.meta.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.