
Real head-to-head · same prompt, one shot
Muse Spark 1.2 vs Fugu Ultra 1.1
Meta's coding reasoning model — co-trained with its own agent, 1M-token window. vs Sakana's multi-agent orchestrator, v1.1 — routes experts per request.
Head-to-head verdict: Fugu Ultra 1.1 wins 11–10 with 2 ties.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Muse Spark 1.2 and Fugu Ultra 1.1, side by side, on 24 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Muse Spark 1.2 · Cloud coder via OpenRouter; the Muse Code agent (one-command install) is its native harness.
Fugu Ultra 1.1 · Benched on GoldieBench via Sakana's Responses API (fugu-ultra-v1.1, xhigh reasoning). Game tasks use the skill-infused threejs-game-director prompt plus a controls+graphics fix loop with an anti-regression clamp, judged on a real mid-play frame by the same Opus judge as the field.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Muse Spark 1.2
Fugu Ultra 1.1
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Page
Visual
Where Muse Spark 1.2 beat Fugu Ultra 1.1
The tasks where I gave Muse Spark 1.2 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Raycaster
Game
Muse Spark 1.2 7.2
·
Fugu Ultra 1.1 2.0
(+5.2)
What I saw: Strong, polished HUD/menu with rich sci-fi framing, minimap, weapon HUD, and a rendering 3D scene visible behind the blurred start overlay — but the screenshot only shows the pre-game menu, so the actual maze walkthrough and enemy/exit gameplay aren't verifiable on-screen, holdin…
Neoncity
Game
Muse Spark 1.2 7.2
·
Fugu Ultra 1.1 2.3
(+4.9)
What I saw: Renders cleanly with a polished HUD (integrity, threat scan, minimap, controls) and a drivable vehicle on a lit road, but the neon-city vibe is weak—buildings read as flat dark boxes with sparse window lights and muted, un-atmospheric lighting rather than the dense glowing cyberp…
Aurora
Visual
Muse Spark 1.2 8.6
·
Fugu Ultra 1.1 5.5
(+3.1)
· Full arctic scene
What I saw: Strong 3D scene with vivid, silky aurora curtains over snowy dunes, silhouetted pine trees, moon and a frozen lake, plus a polished glassy HUD with Kp badge and control chips. Rich composition and glowing shader ray detail push it past the field's best; only minor nit is the some…
Pool
Game
Muse Spark 1.2 7.2
·
Fugu Ultra 1.1 4.5
(+2.7)
What I saw: Renders a clean 3D pool table with properly racked colorful balls, cue ball, pockets, radar minimap, and a polished HUD — clearly on-brief and playable. Weakened by an odd flat teal background/environment, an off-kilter cue prop, and gimmicky 'HP/ENEMIES/haunted' framing that add…
Dogfight
Game
Muse Spark 1.2 8.4
·
Fugu Ultra 1.1 7.2
(+1.2)
What I saw: Strong, polished chase-cam dogfight: crisp HUD with meters/radar/lock indicator, clean 3D plane and volumetric clouds, visible bandits and a lock-on target — very shippable. Slightly held back from top by generic sky-only environment and no visible combat action (tracers/effects)…
Where Fugu Ultra 1.1 beat Muse Spark 1.2
The tasks where I gave Fugu Ultra 1.1 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Outrun
Game
Fugu Ultra 1.1 8.6
·
Muse Spark 1.2 3.2
(+5.4)
· neon combat runner
What I saw: Gorgeous synthwave city with pseudo-3D road, glowing hero craft on a contact disc, and a visible enemy vehicle ahead with combat HUD (KILLS, THREAT HIGH, crosshair) — clearly a combat runner not an empty walking sim. Polished vignette, meters, and banner elevate it above the fiel…
Crypt
Game
Fugu Ultra 1.1 7.8
·
Muse Spark 1.2 4.5
(+3.3)
What I saw: Strong atmospheric torch-lit crypt with a visible skeleton enemy, working HUD (health/torch/kills/depth), coffins and props — clearly a combat crawler not an empty walk sim. Lighting feels a bit washed-out/bright for a 'torch-lit dungeon' and the mood lacks the deep shadowy dread…
Doom
Game
Fugu Ultra 1.1 8.4
·
Muse Spark 1.2 5.5
(+2.9)
· 3D demon shooter
What I saw: Strong atmospheric 3D maze with a visible horned demon enemy, working shotgun, minimap with tracked enemies, and active combat ('CLAWED' hit feedback, HP dropped to 082) — polished HUD and lighting. Falls just shy of the field's best; enemies are more Three.js models than true ra…
Skyrim
Game
Fugu Ultra 1.1 6.8
·
Muse Spark 1.2 4.2
(+2.6)
What I saw: Strong Skyrim-flavored HUD (compass, health/stamina/magicka bars, kills tracker, weapon in hand) and a decent low-poly open world with pines, mountains and dirt path; but the scene looks sparse and the enemy/combat presence is weak — one distant blocky figure and no visible engag…
Dragonflight
Game
Fugu Ultra 1.1 6.5
·
Muse Spark 1.2 4.2
(+2.3)
What I saw: Full polished HUD (health/fury/alt/speed, score, rings, wyverns) and neon rings plus a visible enemy wyvern are present, but the dragon model renders as a jumbled red-blob mess with scattered floating parts, and the scene reads as cluttered/broken rather than a clean flight; comb…
Strengths & weaknesses I logged
Muse Spark 1.2
Strengths
- Generative art & shader-feel scenes (fractal 8.7, aurora/galaxy/matrix/synthwave 8.6)
- Full app chrome one-shot (macOS-clone desktop 8.6)
- Fast one-shots — most builds landed in 45-80s
- 1M context for whole-repo work
Trade-offs
- 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5)
- Open-world briefs collapse to HUD-only shells
- Reasoning tokens billed as output
Fugu Ultra 1.1
Strengths
- Orchestrates 1-3 expert agents per request and synthesises their answers
- Reported SWE-Bench Pro 73.7 — above Opus 4.8 and GPT-5.5 on Sakana's table
- OpenAI- and Anthropic-compatible API — drop-in for Codex and Claude Code
Trade-offs
- Benched as a partial run until the full 50-task batch completes
- Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are
Pricing & context — the spec sheet
| Spec | Muse Spark 1.2 | Fugu Ultra 1.1 |
|---|---|---|
| Vendor | Meta | Sakana AI |
| Context window | 1,000,000 tokens | 1,000,000-token context window |
| Price | $1.25 in / $4.25 out per 1M | API · orchestration billed |
| Pricing detail | Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field. | Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region. |
| Release | 2026-08-05 | 2026-07 |
| Bench coverage | 50/50 scored · avg 7.44/10 | 23/24 scored · avg 6.94/10 |
The verdict — which should you pick?
Across 23 scored shared tasks, the averages are essentially tied — Muse Spark 1.2 7.04 vs Fugu Ultra 1.1 6.94. This isn't the comparison where one wins; it's the comparison where you pick based on context, pricing, and what you're actually trying to ship.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Muse Spark 1.2 and Fugu Ultra 1.1 both into the Agent Operating System and dispatch each from the kanban by task type — generative-art visuals → Muse Spark 1.2, hard, high-stakes coding and reasoning where answer quality beats latency → Fugu Ultra 1.1. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Muse Spark 1.2 vs Fugu Ultra 1.1
Which is better, Muse Spark 1.2 or Fugu Ultra 1.1?
On Goldie Bench, Muse Spark 1.2 averages 7.04/10 across the shared tasks, with 0 gold, 3 silver, 4 bronze overall. Fugu Ultra 1.1 averages 6.94/10, with 0 gold, 1 silver, 2 bronze. Fugu Ultra 1.1 wins the head-to-head 11–10.
How much does Muse Spark 1.2 cost vs Fugu Ultra 1.1?
Muse Spark 1.2: Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field. Fugu Ultra 1.1: Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.
What's the context window for Muse Spark 1.2 vs Fugu Ultra 1.1?
Muse Spark 1.2 has a 1,000,000 tokens context window. Fugu Ultra 1.1 has a 1,000,000-token context window context window.
When should I pick Muse Spark 1.2 over Fugu Ultra 1.1?
Pick Muse Spark 1.2 for: Generative-art visuals; Dashboard & app-shell one-shots; Long-context refactors (1M window). The trade-off is the weaknesses we logged on the bench: 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5); Open-world briefs collapse to HUD-only shells; Reasoning tokens billed as output.
When should I pick Fugu Ultra 1.1 over Muse Spark 1.2?
Pick Fugu Ultra 1.1 for: Hard, high-stakes coding and reasoning where answer quality beats latency; Agentic workflows in Codex / Claude Code via the drop-in provider config; One-shot builds you want a panel of experts on, not a single model. The trade-off is the weaknesses we logged on the bench: Benched as a partial run until the full 50-task batch completes; Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are.
How does Goldie Bench score Muse Spark 1.2 vs Fugu Ultra 1.1?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Muse Spark 1.2 vs Fusion Fugu Ultra 1.1 vs Fusion Muse Spark 1.2 vs Claude Opus 5 Fugu Ultra 1.1 vs Claude Opus 5 Muse Spark 1.2 vs Hermes MoA Fugu Ultra 1.1 vs Hermes MoA Muse Spark 1.2 vs GPT-5.6 Sol Fugu Ultra 1.1 vs GPT-5.6 SolFull model pages: Muse Spark 1.2 · Fugu Ultra 1.1 · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































