Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Real head-to-head · same prompt, one shot

Muse Spark 1.2 vs Fugu Ultra 1.1

Meta's coding reasoning model — co-trained with its own agent, 1M-token window. vs Sakana's multi-agent orchestrator, v1.1 — routes experts per request.

Head-to-head verdict: Fugu Ultra 1.1 wins 11–10 with 2 ties.

Muse Spark 1.2 · context1M tokens
Fugu Ultra 1.1 · context1M tokens
Muse Spark 1.2 · price$1.25 in / $4.25 out per 1M
Fugu Ultra 1.1 · priceAPI · orchestration billed
Muse Spark 1.2 · vendorMeta
Fugu Ultra 1.1 · vendorSakana AI

What I tested — same prompt, two models

I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Muse Spark 1.2 and Fugu Ultra 1.1, side by side, on 24 shared tasks inside the Agent Operating System.

Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.

Muse Spark 1.2 · Cloud coder via OpenRouter; the Muse Code agent (one-command install) is its native harness.

Fugu Ultra 1.1 · Benched on GoldieBench via Sakana's Responses API (fugu-ultra-v1.1, xhigh reasoning). Game tasks use the skill-infused threejs-game-director prompt plus a controls+graphics fix loop with an anti-regression clamp, judged on a real mid-play frame by the same Opus judge as the field.

Side-by-side on 50 shared tasks

Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).

Task ↓
Muse Spark 1.2
Fugu Ultra 1.1
Game
Muse Spark 1.2 on Arcade
Fugu Ultra 1.1 on Arcade
Game
Muse Spark 1.2 on Crypt
Fugu Ultra 1.1 on Crypt
Game
Muse Spark 1.2 on Dogfight
Fugu Ultra 1.1 on Dogfight
Game
Muse Spark 1.2 on Doom
Fugu Ultra 1.1 on Doom
Muse Spark 1.2 on Dragonflight
Fugu Ultra 1.1 on Dragonflight
Muse Spark 1.2 on Dragonrealm
🥉Fugu Ultra 1.1 on Dragonrealm
Game
🥉Muse Spark 1.2 on Flightsim
🥉Fugu Ultra 1.1 on Flightsim
Game
Muse Spark 1.2 on Game
Fugu Ultra 1.1 on Game
Game
Muse Spark 1.2 on Gtadrive
Fugu Ultra 1.1 on Gtadrive
Game
Muse Spark 1.2 on Gtafoot
Fugu Ultra 1.1 on Gtafoot
Muse Spark 1.2 on Neonblaster
Fugu Ultra 1.1 on Neonblaster
Game
Muse Spark 1.2 on Neoncity
Fugu Ultra 1.1 on Neoncity
Game
Muse Spark 1.2 on Neonracer
Fugu Ultra 1.1 on Neonracer
Muse Spark 1.2 on Nordiccrypt
Fugu Ultra 1.1 on Nordiccrypt
Game
Muse Spark 1.2 on Outrun
Fugu Ultra 1.1 on Outrun
Game
Muse Spark 1.2 on Parachute
🥈Fugu Ultra 1.1 on Parachute
Game
Muse Spark 1.2 on Pool
Fugu Ultra 1.1 on Pool
Game
Muse Spark 1.2 on Racing
Fugu Ultra 1.1 on Racing
Game
Muse Spark 1.2 on Raycaster
Fugu Ultra 1.1 on Raycaster
Game
Muse Spark 1.2 on Rpg
Fugu Ultra 1.1 on Rpg
Game
Muse Spark 1.2 on Skyrim
Fugu Ultra 1.1 on Skyrim
Muse Spark 1.2 on Twilightvale
Fugu Ultra 1.1 on Twilightvale
Page
Muse Spark 1.2 on Aipbpromo
Fugu Ultra 1.1 on Aipbpromo
Visual
🥈Muse Spark 1.2 on Aurora
Fugu Ultra 1.1 on Aurora

Where Muse Spark 1.2 beat Fugu Ultra 1.1

The tasks where I gave Muse Spark 1.2 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Raycaster Game
Muse Spark 1.2 7.2 · Fugu Ultra 1.1 2.0 (+5.2)

What I saw: Strong, polished HUD/menu with rich sci-fi framing, minimap, weapon HUD, and a rendering 3D scene visible behind the blurred start overlay — but the screenshot only shows the pre-game menu, so the actual maze walkthrough and enemy/exit gameplay aren't verifiable on-screen, holdin…

Neoncity Game
Muse Spark 1.2 7.2 · Fugu Ultra 1.1 2.3 (+4.9)

What I saw: Renders cleanly with a polished HUD (integrity, threat scan, minimap, controls) and a drivable vehicle on a lit road, but the neon-city vibe is weak—buildings read as flat dark boxes with sparse window lights and muted, un-atmospheric lighting rather than the dense glowing cyberp…

Aurora Visual
Muse Spark 1.2 8.6 · Fugu Ultra 1.1 5.5 (+3.1) · Full arctic scene

What I saw: Strong 3D scene with vivid, silky aurora curtains over snowy dunes, silhouetted pine trees, moon and a frozen lake, plus a polished glassy HUD with Kp badge and control chips. Rich composition and glowing shader ray detail push it past the field's best; only minor nit is the some…

Pool Game
Muse Spark 1.2 7.2 · Fugu Ultra 1.1 4.5 (+2.7)

What I saw: Renders a clean 3D pool table with properly racked colorful balls, cue ball, pockets, radar minimap, and a polished HUD — clearly on-brief and playable. Weakened by an odd flat teal background/environment, an off-kilter cue prop, and gimmicky 'HP/ENEMIES/haunted' framing that add…

Dogfight Game
Muse Spark 1.2 8.4 · Fugu Ultra 1.1 7.2 (+1.2)

What I saw: Strong, polished chase-cam dogfight: crisp HUD with meters/radar/lock indicator, clean 3D plane and volumetric clouds, visible bandits and a lock-on target — very shippable. Slightly held back from top by generic sky-only environment and no visible combat action (tracers/effects)…

Where Fugu Ultra 1.1 beat Muse Spark 1.2

The tasks where I gave Fugu Ultra 1.1 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Outrun Game
Fugu Ultra 1.1 8.6 · Muse Spark 1.2 3.2 (+5.4) · neon combat runner

What I saw: Gorgeous synthwave city with pseudo-3D road, glowing hero craft on a contact disc, and a visible enemy vehicle ahead with combat HUD (KILLS, THREAT HIGH, crosshair) — clearly a combat runner not an empty walking sim. Polished vignette, meters, and banner elevate it above the fiel…

Crypt Game
Fugu Ultra 1.1 7.8 · Muse Spark 1.2 4.5 (+3.3)

What I saw: Strong atmospheric torch-lit crypt with a visible skeleton enemy, working HUD (health/torch/kills/depth), coffins and props — clearly a combat crawler not an empty walk sim. Lighting feels a bit washed-out/bright for a 'torch-lit dungeon' and the mood lacks the deep shadowy dread…

Doom Game
Fugu Ultra 1.1 8.4 · Muse Spark 1.2 5.5 (+2.9) · 3D demon shooter

What I saw: Strong atmospheric 3D maze with a visible horned demon enemy, working shotgun, minimap with tracked enemies, and active combat ('CLAWED' hit feedback, HP dropped to 082) — polished HUD and lighting. Falls just shy of the field's best; enemies are more Three.js models than true ra…

Skyrim Game
Fugu Ultra 1.1 6.8 · Muse Spark 1.2 4.2 (+2.6)

What I saw: Strong Skyrim-flavored HUD (compass, health/stamina/magicka bars, kills tracker, weapon in hand) and a decent low-poly open world with pines, mountains and dirt path; but the scene looks sparse and the enemy/combat presence is weak — one distant blocky figure and no visible engag…

Fugu Ultra 1.1 6.5 · Muse Spark 1.2 4.2 (+2.3)

What I saw: Full polished HUD (health/fury/alt/speed, score, rings, wyverns) and neon rings plus a visible enemy wyvern are present, but the dragon model renders as a jumbled red-blob mess with scattered floating parts, and the scene reads as cluttered/broken rather than a clean flight; comb…

Strengths & weaknesses I logged

Muse Spark 1.2

Strengths

  • Generative art & shader-feel scenes (fractal 8.7, aurora/galaxy/matrix/synthwave 8.6)
  • Full app chrome one-shot (macOS-clone desktop 8.6)
  • Fast one-shots — most builds landed in 45-80s
  • 1M context for whole-repo work

Trade-offs

  • 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5)
  • Open-world briefs collapse to HUD-only shells
  • Reasoning tokens billed as output

Fugu Ultra 1.1

Strengths

  • Orchestrates 1-3 expert agents per request and synthesises their answers
  • Reported SWE-Bench Pro 73.7 — above Opus 4.8 and GPT-5.5 on Sakana's table
  • OpenAI- and Anthropic-compatible API — drop-in for Codex and Claude Code

Trade-offs

  • Benched as a partial run until the full 50-task batch completes
  • Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are

Pricing & context — the spec sheet

Spec Muse Spark 1.2 Fugu Ultra 1.1
VendorMetaSakana AI
Context window1,000,000 tokens1,000,000-token context window
Price$1.25 in / $4.25 out per 1MAPI · orchestration billed
Pricing detailMeta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field.Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.
Release2026-08-052026-07
Bench coverage50/50 scored · avg 7.44/1023/24 scored · avg 6.94/10

The verdict — which should you pick?

Across 23 scored shared tasks, the averages are essentially tied — Muse Spark 1.2 7.04 vs Fugu Ultra 1.1 6.94. This isn't the comparison where one wins; it's the comparison where you pick based on context, pricing, and what you're actually trying to ship.

If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Muse Spark 1.2 and Fugu Ultra 1.1 both into the Agent Operating System and dispatch each from the kanban by task type — generative-art visuals → Muse Spark 1.2, hard, high-stakes coding and reasoning where answer quality beats latency → Fugu Ultra 1.1. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.

FAQ — Muse Spark 1.2 vs Fugu Ultra 1.1

Which is better, Muse Spark 1.2 or Fugu Ultra 1.1?

On Goldie Bench, Muse Spark 1.2 averages 7.04/10 across the shared tasks, with 0 gold, 3 silver, 4 bronze overall. Fugu Ultra 1.1 averages 6.94/10, with 0 gold, 1 silver, 2 bronze. Fugu Ultra 1.1 wins the head-to-head 11–10.

How much does Muse Spark 1.2 cost vs Fugu Ultra 1.1?

Muse Spark 1.2: Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field. Fugu Ultra 1.1: Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.

What's the context window for Muse Spark 1.2 vs Fugu Ultra 1.1?

Muse Spark 1.2 has a 1,000,000 tokens context window. Fugu Ultra 1.1 has a 1,000,000-token context window context window.

When should I pick Muse Spark 1.2 over Fugu Ultra 1.1?

Pick Muse Spark 1.2 for: Generative-art visuals; Dashboard & app-shell one-shots; Long-context refactors (1M window). The trade-off is the weaknesses we logged on the bench: 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5); Open-world briefs collapse to HUD-only shells; Reasoning tokens billed as output.

When should I pick Fugu Ultra 1.1 over Muse Spark 1.2?

Pick Fugu Ultra 1.1 for: Hard, high-stakes coding and reasoning where answer quality beats latency; Agentic workflows in Codex / Claude Code via the drop-in provider config; One-shot builds you want a panel of experts on, not a single model. The trade-off is the weaknesses we logged on the bench: Benched as a partial run until the full 50-task batch completes; Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are.

How does Goldie Bench score Muse Spark 1.2 vs Fugu Ultra 1.1?

Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly