Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Sakana AI

Fugu Ultra

Sakana's multi-agent answer to Fusion — frontier ensemble without single-vendor risk.

Context272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates.
Pricing$5 / 1M input · $30 / 1M output (Fugu Ultra)
Tasks tested42
Avg score7.94/10 average
Medals🥇5 🥈2 🥉2
Release2026-06-15
Official sitesakana.ai ↗
Official vendor source
Fugu Ultra is built by Sakana AI — see the vendor's own product page, pricing, and docs at sakana.ai.
Visit sakana.ai →

Reference benchmarks for Fugu Ultra

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Fugu Ultra is honest about what's measured.

SWE Bench Pro
73.7
GPQA-Diamond
95.5
MRCRv2
93.6

What is Fugu Ultra?

Fugu Ultra is the Sakana AI frontier model with a 272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates. context window, released 2026-06-15. Tagline: Sakana's multi-agent answer to Fusion — frontier ensemble without single-vendor risk.. Official source: sakana.ai.

Pricing detail. Sakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M output for Fugu Ultra. Direct competitor to OpenRouter Fusion's panel approach.

How I use it inside the Agent OS. Dispatched from Agent OS as the panel-ensemble alternative to OpenRouter Fusion. Bench scored by Claude judge against the same 42 prompts as every other model.

What I built with Fugu Ultra

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Fugu Ultra shipped on the bench: 42 one-shot demos across 272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates. of context. Of those, 42 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • SWE Bench Pro 73.7 · GPQA-D 95.5 · MRCRv2 93.6 — Sakana's published frontier-tier benchmark scores
  • Vendor-agnostic ensemble — opt out of specific providers for compliance / export-control
  • OpenAI-compatible API at api.sakana.ai — drop-in for existing tooling

Trade-offs

  • Panel orchestration adds latency — even a 'pong' burns ~2k orchestration tokens
  • Newer than Fusion; less community calibration on long-tail prompts

Best for

  • Teams that want Fusion-class quality but need a different vendor risk profile
  • Operators avoiding export-controlled providers (Sakana emphasises this in their pitch)
  • Deep-research workflows where ensemble verdicts beat single-model answers

Every demo by Fugu Ultra

42 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Fugu Ultra one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
Ultra v2 (gap-fill) — first-person Nordic dungeon. Smoke-test MAYBE (0.4% diff) — pointer-lock FPS; flagged for manual verification.
Fugu Ultra one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
Ultra v2 — Skyrim-style frozen open world. Smoke-test MAYBE (0.1% diff) — loads, minimal change on generic input.
Fugu Ultra one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
Ultra v2 (gap-fill, 16-min direct call) — 45KB Doom-style raycaster FPS: sprite enemies, gun + muzzle flash, ammo/health HUD, 2 rAF loops, 9 input handlers. Smoke-test STATIC because movement is gated on pointer-lock (update() returns while the start-overlay is open) and headless Chrome can't engage pointer-lock — same limitation as skyrim/crypt. Renders a full 274KB scene; flagged for manual verification.
Fugu Ultra one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
Ultra v2 — 3D arcade racer with drift. Smoke-test PASS.
Fugu Ultra one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
Ultra v2 (gap-fill) — Skyrim-style open world. Smoke-test MAYBE (0.0% diff) — FPS mouse-look needs pointer-lock which generic input didn't engage; flagged for manual verification.
Fugu Ultra one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale 🥉
Game
Ultra v2 — 61.8KB open-world RPG (village, NPCs, weather, day/night). Smoke-test PASS. Densest Ultra build on the bench.
Fugu Ultra one-shot build of Voxelcraft — GoldieBench AI benchmark screenshot▶ LIVE
Voxelcraft
Game
Ultra v2 (gap-fill) — Minecraft-style voxel sandbox. Smoke-test PASS.
Fugu Ultra one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade
Game
Ultra v2 — Breakout-style. Smoke-test MAYBE (0.4% diff) — runs but generic click+WASD didn't trigger strong motion; paddle may need mouse-move.
Fugu Ultra one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
Ultra v2 (gap-fill) — 3D dogfight. Smoke-test MAYBE (0.0% diff) — flight controls may need specific keys; flagged for manual verification.
Fugu Ultra one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight 🥉
Game
Ultra v2 — dragon through neon rings, full HUD. Smoke-test PASS (4.2% pixel diff).
Fugu Ultra one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game 🥇
Game
Ultra v2 — juicy browser game. Smoke-test PASS with 24.7% pixel diff — one of the most reactive builds on the bench.
Fugu Ultra one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
Ultra v2 — arcade shooter. Smoke-test MAYBE (0.0% diff) — may need a specific start interaction; flagged for manual verification.
Fugu Ultra one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity
Game
Ultra v2 — cyberpunk neon-city flythrough. Smoke-test PASS (9.1% pixel diff).
Fugu Ultra one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
Ultra v2 — top-down neon racer. Smoke-test PASS.
Fugu Ultra one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt 🥈
Game
Ultra v2 (gap-fill) — 61.5KB Nordic dungeon crawler with bloom + boss room. Smoke-test PASS with 22.8% pixel diff — highly reactive.
Fugu Ultra one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
Ultra v2 (gap-fill, 19-min direct call) — 36KB pseudo-3D OutRun racer: curving/cresting road, synthwave sunset, parallax mountains, 2 rAF loops. Smoke-test MAYBE (0.1% diff) — road advances on held-accelerate which the brief keypress didn't sustain; flagged for manual verification.
Fugu Ultra one-shot build of Pool — GoldieBench AI benchmark screenshot▶ LIVE
Pool
Game
Ultra v2 — billiards. Smoke-test MAYBE (0.0% diff) — cue needs click-drag aim that generic input didn't replicate.
Fugu Ultra one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster 🥈
Game
26KB canvas raycaster with WASD + mouse-look + distance fog + weapon bob. Clean implementation, comparable to Fusion's 17KB on the same prompt. ~$0.35 per call — roughly 1/4 the cost of Fusion.
Fugu Ultra one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
Ultra v2 (gap-fill) — top-down RPG with tilemap + NPCs. Smoke-test PASS (1.1% diff).
Fugu Ultra one-shot build of Landing — GoldieBench AI benchmark screenshot▶ LIVE
Landing 🥇
Page
Sakana Fugu Ultra shipped a 32KB Apple-keynote landing — bigger than Fusion's 20KB attempt at the same prompt. Animated mesh gradient, multi-section, polished. $0.32 vs Fusion's $1.30 for the same output — 4× cheaper, denser result.
Fugu Ultra one-shot build of Webos — GoldieBench AI benchmark screenshot▶ LIVE
Webos
Page
Ultra v2 — desktop OS. Smoke-test MAYBE (0.3% diff) — static desktop until you open an app; expected for a UI-shell build.
Fugu Ultra one-shot build of Blackhole — GoldieBench AI benchmark screenshot▶ LIVE
Blackhole
Sim
Ultra v2 — gravitational-lensing black hole. Smoke-test PASS (2.2% pixel diff).
Fugu Ultra one-shot build of Boids — GoldieBench AI benchmark screenshot▶ LIVE
Boids
Sim
Ultra v2 — boids flock, agents respond to mouse. Smoke-test PASS (9.5% pixel diff — strong motion).
Fugu Ultra one-shot build of Cloth — GoldieBench AI benchmark screenshot▶ LIVE
Cloth
Sim
Ultra v2 — verlet cloth sim, drag to deform. Smoke-test PASS.
Fugu Ultra one-shot build of Fluid — GoldieBench AI benchmark screenshot▶ LIVE
Fluid
Sim
Ultra v2 — fluid sim. Smoke-test MAYBE (0.0% diff) — needs click-drag injection which generic input didn't replicate.
Fugu Ultra one-shot build of Fractal — GoldieBench AI benchmark screenshot▶ LIVE
Fractal
Sim
Ultra v2 — Mandelbrot zoom. Smoke-test PASS.
Fugu Ultra one-shot build of Galaxy — GoldieBench AI benchmark screenshot▶ LIVE
Galaxy
Sim
26KB three.js spiral galaxy with drag-to-orbit + dust lanes + bloom. Comparable visual quality to Fusion's 14KB attempt with more polish on the camera UI. ~$0.24 per call.
Fugu Ultra one-shot build of Orbit — GoldieBench AI benchmark screenshot▶ LIVE
Orbit
Sim
26KB inner-solar-system orbit map with a glassmorphic info panel, kicker badge, blurred backdrop, hover cards. Cleaner UI than Fusion's same-task attempt — beats it on polish.
Fugu Ultra one-shot build of Particleforge — GoldieBench AI benchmark screenshot▶ LIVE
Particleforge
Sim
Ultra v2 (gap-fill) — mouse-gravity particle sculptor. Smoke-test PASS (2.6% diff).
Fugu Ultra one-shot build of Pathtracer — GoldieBench AI benchmark screenshot▶ LIVE
Pathtracer
Sim
Ultra v2 — WebGL path tracer with sample accumulation. Smoke-test PASS (4.1% pixel diff).
Fugu Ultra one-shot build of Reactiondiff — GoldieBench AI benchmark screenshot▶ LIVE
Reactiondiff
Sim
Ultra v2 — Gray-Scott reaction-diffusion. Smoke-test PASS.
Fugu Ultra one-shot build of Solar — GoldieBench AI benchmark screenshot▶ LIVE
Solar 🥇
Sim
Fugu Ultra v2 rebuild — 55.7KB solar system, the densest solar attempt on the bench. Full </html>, animation loop, Saturn rings, drag-to-orbit + scroll-to-zoom. Smoke-test PASS (3.6% pixel diff after drag, zero console errors). The panel ensemble produces a markedly richer build than Mini's 12KB single-model version.
Fugu Ultra one-shot build of Wormhole — GoldieBench AI benchmark screenshot▶ LIVE
Wormhole
Sim
Ultra v2 — wormhole tunnel. Smoke-test MAYBE (0.3% diff) — animates slowly; hold-space accel didn't show big change in the short test.
Fugu Ultra one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora
Visual
Ultra v2 (gap-fill) — aurora over mountain ridge. Smoke-test PASS (0.8% diff — visual-only prompt).
Fugu Ultra one-shot build of Fireworks — GoldieBench AI benchmark screenshot▶ LIVE
Fireworks 🥇
Visual
Ultra v2 (gap-fill) — click-to-launch fireworks. Smoke-test PASS with 26.3% pixel diff — among the most reactive builds.
Fugu Ultra one-shot build of Lavalamp — GoldieBench AI benchmark screenshot▶ LIVE
Lavalamp
Visual
Ultra v2 — lava lamp. Smoke-test MAYBE (0.1% diff) — slow ambient motion is expected for this visual-only prompt.
Fugu Ultra one-shot build of Matrix — GoldieBench AI benchmark screenshot▶ LIVE
Matrix
Visual
Ultra v2 — Matrix rain. Smoke-test PASS (4.6%).
Fugu Ultra one-shot build of Plasma — GoldieBench AI benchmark screenshot▶ LIVE
Plasma
Visual
Ultra v2 — plasma effect with palette switcher. Smoke-test PASS.
Fugu Ultra one-shot build of Synthwave — GoldieBench AI benchmark screenshot▶ LIVE
Synthwave
Visual
Ultra v2 — synthwave terrain flythrough. Smoke-test PASS (4.3% pixel diff).
Fugu Ultra one-shot build of Terrain — GoldieBench AI benchmark screenshot▶ LIVE
Terrain
Visual
Ultra v2 — Tron procedural terrain. Smoke-test PASS.
Fugu Ultra one-shot build of Voxel — GoldieBench AI benchmark screenshot▶ LIVE
Voxel 🥇
Visual
Ultra v2 — Temple-Run voxel runner. Smoke-test PASS with 32.7% pixel diff — the single most reactive build. Replaces the earlier truncated voxel-fugu that was deleted.
Fugu Ultra one-shot build of Waves — GoldieBench AI benchmark screenshot▶ LIVE
Waves
Visual
Ultra v2 — Gerstner ocean waves. Smoke-test PASS (3.6% pixel diff).
every demo, in a grid · click any one to play

Compare Fugu Ultra against every other model

Every head-to-head featuring Fugu Ultra. Verdicts shown for scored pairs.

Fugu Ultra vs Fusion
Fusion leads 26–4
Fugu Ultra vs Claude Opus 5
Claude Opus 5 leads 28–13
Fugu Ultra vs Hermes MoA
Hermes MoA leads 27–15
Fugu Ultra vs GPT-5.6 Sol
GPT-5.6 Sol leads 28–13
Fugu Ultra vs Claude Fable 5
Claude Fable 5 leads 25–16
Fugu Ultra vs Qwen 3.8
Qwen 3.8 leads 26–11
Fugu Ultra vs Grok
Fugu Ultra leads 16–14
Fugu Ultra vs MiniMax M3
MiniMax M3 leads 17–15
Fugu Ultra vs Kimi K3
Kimi K3 leads 27–15
Fugu Ultra vs GLM-5.2
Fugu Ultra leads 23–15
Fugu Ultra vs Fugu Mini
Fugu Ultra leads 18–9
Fugu Ultra vs Muse Spark 1.2
Fugu Ultra leads 23–19
Fugu Ultra vs Opus 4.8
Fugu Ultra leads 25–11
Fugu Ultra vs Kimi K2.7
Fugu Ultra leads 12–6
Fugu Ultra vs Qwable 5 27B Coder
Fugu Ultra leads 28–12
Fugu Ultra vs Gemini 3.6 Flash
Fugu Ultra leads 28–14
Fugu Ultra vs Claude Sonnet 5
Fugu Ultra leads 25–15
Fugu Ultra vs Qwen 3.7
Fugu Ultra leads 34–6
Fugu Ultra vs Fugu Ultra 1.1
Fugu Ultra leads 10–8
Fugu Ultra vs Inkling
Fugu Ultra leads 37–5
Fugu Ultra vs Agents-A1
Fugu Ultra leads 39–3
Fugu Ultra vs Gemma 4 12B · MLX
Fugu Ultra leads 41–0
Fugu Ultra vs Laguna XS 2.1
Fugu Ultra leads 42–0
Fugu Ultra vs Qwythos 9B
Fugu Ultra leads 42–0
Fugu Ultra vs LongCat-2.0
LongCat-2.0 leads 3–1
Fugu Ultra vs Hy3
Tied 1–1
Fugu Ultra vs Gemma-4 12B Coder
Fugu Ultra leads 6–0
Fugu Ultra vs DeepSeek V4 Flash
42 shared tasks · unscored
Fugu Ultra vs DeepSeek V4 Pro
42 shared tasks · unscored
Fugu Ultra vs Kimi K2.7 · Fast
42 shared tasks · unscored
Fugu Ultra vs Kimi K2.7 · No-Think
42 shared tasks · unscored
Fugu Ultra vs Kimi K2.7 · Quality
42 shared tasks · unscored
Fugu Ultra vs Ornith 1.0
42 shared tasks · unscored
Fugu Ultra vs Claude Mythos 5
Reference-only
Fugu Ultra vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Fugu Ultra vs Fusion Fugu Ultra vs Claude Opus 5 Fugu Ultra vs Hermes MoA Fugu Ultra vs GPT-5.6 Sol Fugu Ultra vs Claude Fable 5 Fugu Ultra vs Qwen 3.8 Fugu Ultra vs Grok Fugu Ultra vs MiniMax M3 Fugu Ultra vs Kimi K3 Fugu Ultra vs GLM-5.2 Fugu Ultra vs Fugu Mini Fugu Ultra vs Muse Spark 1.2 Fugu Ultra vs Opus 4.8 Fugu Ultra vs Kimi K2.7 Fugu Ultra vs Qwable 5 27B Coder Fugu Ultra vs Gemini 3.6 Flash Fugu Ultra vs Claude Sonnet 5 Fugu Ultra vs Qwen 3.7 Fugu Ultra vs Fugu Ultra 1.1 Fugu Ultra vs Inkling Fugu Ultra vs Agents-A1 Fugu Ultra vs Gemma 4 12B · MLX Fugu Ultra vs Laguna XS 2.1 Fugu Ultra vs Qwythos 9B Fugu Ultra vs LongCat-2.0 Fugu Ultra vs Hy3 Fugu Ultra vs Gemma-4 12B Coder

Read more on agentos.guide: /sakana-fugu-vs-fusion

Fugu Ultra — frequently asked

What is Fugu Ultra?

Fugu Ultra is Sakana AI's AI model — Sakana's multi-agent answer to Fusion — frontier ensemble without single-vendor risk. It has a 272K tokens (free) · larger via paid tier context window and was released 2026-06-15.

How good is Fugu Ultra at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 7.94/10 across 42 scored tasks, with 5 gold, 2 silver and 2 bronze medals.

How much does Fugu Ultra cost?

$5 / 1M input · $30 / 1M output (Fugu Ultra). Sakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M ou

Where can I see Fugu Ultra demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly