Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Muse Spark 1.2's full benchmark breakdown.

Meta's coding reasoning model — co-trained with its own agent, 1M-token window. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-08-06
Scored tasks
50
Reading time
10 min
External sources
1

01 · The headline numbers

GoldieBench average
7.47/10
50 scored one-shot tasks
Board rank
#15
of 22 ranked frontier models
Task medals
0🥇 3🥈 4🥉
outright wins on shared briefs
Context
1,000,000 tokens
$1.25 in / $4.25 out per 1M

02 · Every benchmark, bar by bar

All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Strong: a beautifully rendered GPU Julia set with smooth coloring, orbit-trap glow, and a polished glassmorphic control panel (mode toggle, iterations, palettes, live zoom/center stats). Weak: nothing major visible — a very complete, on-brief, task-winning build.”

- the judge's verdict, unedited

“Strong 3D scene with vivid, silky aurora curtains over snowy dunes, silhouetted pine trees, moon and a frozen lake, plus a polished glassy HUD with Kp badge and control chips. Rich composition and glowing shader ray detail push it past the field's best; only minor nit is the somewhat flat foreground snow expanse.”

- the judge's verdict, unedited

“Gorgeous dense rain with bright white heads, proper katakana glyphs, and a cohesive cyberpunk HUD (stream status, carrier signal, control bar) that elevates it well past a generic canvas demo. Strong glow, vignette, and interactivity hooks make it a task winner; the only minor nit is the 3D Three.js layer being mostly hidden behind the dominant rain.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Muse Spark 1.2's lowest scored builds, verdicts unedited:

“The HUD renders cleanly with polished synthwave styling, but the entire 3D scene — road, sun, mountains, car — is a black void, so the actual outrun driving game never appears. A functional shell around a non-rendering canvas fails the core brief.”

- the judge's verdict, unedited

“The HUD is polished and thematically on-brief (compass, health/stamina/magicka bars, Cinzel fonts), but the 3D world is bare and unconvincing—flat washed-out terrain, only a couple of low-poly trees and rocks, and a giant floating gray monolith that looks like a broken mesh. No visible enemies, buildings, or sense of an explorable Skyrim world, making it feel unfinished versus the field's best.”

- the judge's verdict, unedited

“HUD is genuinely polished — health/fury/speed meters, altitude/rings/combo, score+best, minimap, and control hints all look shippable — but the actual 3D scene is broken: the camera is buried inside a giant cone/horn with no visible dragon body, no neon rings, no canyon, and no sky. The core gameplay view fails to render on-brief.”

- the judge's verdict, unedited

“The HUD is polished and on-brief (torch fuel, vitalis, kills, coffin altar), but the rendered scene looks flat and bright rather than torch-lit and dark — lighting is washed-out with pale beige walls, floating spheres, and no atmospheric torch glow, badly undermining the dungeon-crawler mood.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (23 tasks)

Muse Spark 1.27.0
field avg7.2

Others (3 tasks)

Muse Spark 1.27.1
field avg7.5

Pages (3 tasks)

Muse Spark 1.28.2
field avg7.6

Sims (12 tasks)

Muse Spark 1.28.2
field avg6.9

Visuals (9 tasks)

Muse Spark 1.27.5
field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Muse Spark 1.2 perform on GoldieBench?

Muse Spark 1.2 averages 7.47/10 across 50 scored one-shot build tasks, ranking #15 of 22 ranked frontier models, with 0 task golds, 3 silvers and 4 bronzes. Every score is a real judged build you can open and play on this site.

02What is Muse Spark 1.2 best at?

Its strongest scored build is Fractal at 8.7/10. By category it averages Game 7.0, Other 7.1, Page 8.2, Sim 8.2, Visual 7.5 on the bench.

03How much does Muse Spark 1.2 cost?

Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field.

Source ledger

  1. 01Official vendor site: developer.meta.com/ai/models/muse-sparkdeveloper.meta.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly