Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Sakana AI

Fugu Ultra 1.1

Sakana's multi-agent orchestrator, v1.1 — routes experts per request.

Context1,000,000-token context window
PricingAPI · orchestration billed
Tasks tested24
Avg score6.94/10 average
Medals🥇0 🥈1 🥉2
Release2026-07
Official sitesakana.ai ↗
Official vendor source
Fugu Ultra 1.1 is built by Sakana AI — see the vendor's own product page, pricing, and docs at sakana.ai.
Visit sakana.ai →

Reference benchmarks for Fugu Ultra 1.1

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Fugu Ultra 1.1 is honest about what's measured.

SWE-Bench Pro (reported)
73.7
source: /sakana.ai
LiveCodeBench (reported)
93.2
source: /sakana.ai
GPQA Diamond (reported)
95.5
source: /sakana.ai

What is Fugu Ultra 1.1?

Fugu Ultra 1.1 is the Sakana AI frontier model with a 1,000,000-token context window context window, released 2026-07. Tagline: Sakana's multi-agent orchestrator, v1.1 — routes experts per request.. Official source: sakana.ai.

Pricing detail. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.

How I use it inside the Agent OS. Benched on GoldieBench via Sakana's Responses API (fugu-ultra-v1.1, xhigh reasoning). Game tasks use the skill-infused threejs-game-director prompt plus a controls+graphics fix loop with an anti-regression clamp, judged on a real mid-play frame by the same Opus judge as the field.

What I built with Fugu Ultra 1.1

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Fugu Ultra 1.1 shipped on the bench: 24 one-shot demos across 1,000,000-token context window of context. Of those, 23 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Orchestrates 1-3 expert agents per request and synthesises their answers
  • Reported SWE-Bench Pro 73.7 — above Opus 4.8 and GPT-5.5 on Sakana's table
  • OpenAI- and Anthropic-compatible API — drop-in for Codex and Claude Code

Trade-offs

  • Benched as a partial run until the full 50-task batch completes
  • Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are

Best for

  • Hard, high-stakes coding and reasoning where answer quality beats latency
  • Agentic workflows in Codex / Claude Code via the drop-in provider config
  • One-shot builds you want a panel of experts on, not a single model

Every demo by Fugu Ultra 1.1

24 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Fugu Ultra 1.1 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
Strong atmospheric torch-lit crypt with a visible skeleton enemy, working HUD (health/torch/kills/depth), coffins and props — clearly a combat crawler not an empty walk sim. Lighting feels a bit washed-out/bright for a 'torch-lit dungeon' and the mood lacks the deep shadowy dread of the top entry, keeping it shippable but not quite a task winner.
Fugu Ultra 1.1 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm 🥉
Game
Strong Skyrim-vibe frozen open world with layered snowy mountains, a player with visible sword, multiple approaching enemies with health orbs, runes, an event banner ('DRAGON SWOOP · FIRE BREATH') and polished HUD/compass — clearly beats the empty-walking-sim trap. Enemy models are a touch simplistic and the moon/sky are basic, but the visible combat and worldbuilding make it a task winner.
Fugu Ultra 1.1 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
Strong atmospheric 3D maze with a visible horned demon enemy, working shotgun, minimap with tracked enemies, and active combat ('CLAWED' hit feedback, HP dropped to 082) — polished HUD and lighting. Falls just shy of the field's best; enemies are more Three.js models than true raycaster sprites, but combat clearly works.
Fugu Ultra 1.1 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
Gorgeous polished third-person racer with a clear track, guardrails, trees/rocks/buildings, obstacle cones, a slick craft with ground shadow, and combat layered in (crosshair, KILLS 0/12, visible enemies/targets ahead) plus rich HUD with minimap — strong shippable build; only minor knock is kills still at 00 in this frame, but the enemy/combat system is clearly present and working.
Fugu Ultra 1.1 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
Strong Skyrim-flavored HUD (compass, health/stamina/magicka bars, kills tracker, weapon in hand) and a decent low-poly open world with pines, mountains and dirt path; but the scene looks sparse and the enemy/combat presence is weak — one distant blocky figure and no visible engaged fight, so it reads more explorer than the best in field.
Fugu Ultra 1.1 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
Gorgeous cohesive twilight scene with detailed hero holding a weapon, visible enemies (creature + humanoid), wooden bridge, lampposts, cottage, river, pickups and a working minimap/HUD with quest and kill tracker. Strong art direction and populated world with combat framing; only minor risk is verifying enemy AI aggression, but the visible enemies plus polish put it at the top of the field.
Fugu Ultra 1.1 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
Polished HUD (health/ammo/wanted stars/minimap with tracked entities) and a well-modeled character with active shooting (ammo already at 089, 'CIVILIANS SCATTER' banner), but the camera is clipped hard into a building wall showing mostly empty geometry and no visible enemies/combat in frame, so it reads more like a functional walking sim than the shootout the brief demands.
Fugu Ultra 1.1 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
Strong polished HUD with wanted stars, minimap showing cops/traffic dots, and combat state ('UNDER FIRE', health at 74 taking damage) plus debris particles indicate working enemies; weakened by the player car's messy/broken 3D model that looks like clipped blocks rather than a clean vehicle.
Fugu Ultra 1.1 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo
Page
Fugu Ultra 1.1 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute 🥈
Game
Strong deployed-chute skydiver over a jungle canopy with a polished HUD (altitude, dist-to-H, score) plus active drone enemies, flare combat, and 'THREAT DOWN' kill feedback — it delivers the full jump/steer/land loop AND working combat, edging past a bland walking sim.
Fugu Ultra 1.1 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim 🥉
Game
Polished flight sim with a detailed aircraft model, full HUD (airspeed/alt/VS/heading tape/attitude indicator), runway with markings, hangar, control tower and terrain — plus a combat layer with visible drones and cannon reticle. Very strong and shippable, but the enemy at this frame is just a single distant drone and it sits right at the top of the field's best rather than clearly beating it.
Fugu Ultra 1.1 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade
Game
The brief asked for a classic arcade game (tetris/breakout/snake); this is a 3D flying-shooter titled 'Breakout Hunter' with drones and radar — polished HUD and clean low-poly visuals, but it completely misses the arcade brief and there's no actual breakout/paddle/brick gameplay visible despite the label. Strong presentation, wrong genre.
Fugu Ultra 1.1 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
Strong visuals — detailed hero jet with shadow, polished HUD, city skyline and terrain create a real air-combat atmosphere; but the screenshot shows no visible enemies or combat (0/10 kills, no dogfight action), so it reads closer to a flight-sim showcase than a proven shooter.
Fugu Ultra 1.1 one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight
Game
Full polished HUD (health/fury/alt/speed, score, rings, wyverns) and neon rings plus a visible enemy wyvern are present, but the dragon model renders as a jumbled red-blob mess with scattered floating parts, and the scene reads as cluttered/broken rather than a clean flight; combat and fire-breath appear plausible in source but the visual execution undercuts it.
Fugu Ultra 1.1 one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game
Game
Polished 3D low-poly arena shooter with clean HUD, working radar showing 9 hostiles, HP damage ('ARMOR HIT', HP down to 045) and full combat controls—clearly a functional game, not a walking sim. Held back from top tier by the odd/janky player craft model and no enemies actually visible on-screen in this frame, so combat readability is weaker than the radar suggests.
Fugu Ultra 1.1 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
Polished 3D low-poly arcade shooter with visible enemies, active combat state (score 238, kills 1, health 75, 5 enemies), power-up pickups, crystals/asteroids scenery and a clean neon HUD — clearly functional and juicy; slightly generic ground-based feel rather than a true space shooter and boss/screen-shake not evident in the still.
Fugu Ultra 1.1 one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity
Game
HUD panels (health/boost/kills/crosshair) render nicely but the entire 3D scene is black — no city, road, car, or enemies visible, indicating the Three.js world failed to render. A non-functional walking/driving sim regardless of the ambitious source.
Fugu Ultra 1.1 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
Strong neon vaporwave scene with detailed craft, gradient sky, retro sun stripes, and a polished HUD showing 9 enemies to hunt; the vapor-trail particle effects aren't visible in this static shot and no enemies are yet in frame, keeping it just shy of the field's best.
Fugu Ultra 1.1 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
Strong HUD, first-person weapon/shield and visible enemies (7) with combat scaffolding are present, but the scene reads bright and flat rather than torch-lit crypt — lighting/atmosphere totally missing the Nordic dungeon mood, and floating stretched geometry looks buggy.
Fugu Ultra 1.1 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
Gorgeous synthwave city with pseudo-3D road, glowing hero craft on a contact disc, and a visible enemy vehicle ahead with combat HUD (KILLS, THREAT HIGH, crosshair) — clearly a combat runner not an empty walking sim. Polished vignette, meters, and banner elevate it above the field's best; only minor knock is the driving game reads more as a hover-shooter than pure Outrun cruiser.
Fugu Ultra 1.1 one-shot build of Pool — GoldieBench AI benchmark screenshot▶ LIVE
Pool
Game
Visually polished 3D pool-hall arena with a giant table, pockets, minimap and combat HUD, but this is an arena shooter reskin ('cue-knight', enemies, HP/boost) — not a physically simulated billiards game as the brief demands, so it misses the core task entirely despite decent rendering.
Fugu Ultra 1.1 one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster
Game
The HUD (health, charge, minimap frame, controls) renders but the entire 3D scene is black with no visible maze, walls, enemies, or floor — the raycaster world failed to render. HOSTILES 0/0 and empty minimap confirm the core build is broken/non-rendering.
Fugu Ultra 1.1 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
Polished top-down 3D RPG with visible enemies in aggro rings, active combat ('HIT -13' damage numbers, health at 87), pickups, a shrine objective, and a clean HUD showing inventory/kills/hostiles; strong shippable build, only slightly held back by the inventory system being minimal rather than a true managed bag.
Fugu Ultra 1.1 one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora
Visual
Clean 3D scene with moon, starfield, layered mountains and pyramid 'trees', but the actual aurora is barely visible—only faint green ground-glow blobs rather than the rippling northern-lights sky the brief demands. Polished framing and typography, but it fails the core aurora effect.
every demo, in a grid · click any one to play

Compare Fugu Ultra 1.1 against every other model

Every head-to-head featuring Fugu Ultra 1.1. Verdicts shown for scored pairs.

Fugu Ultra 1.1 vs Fusion
Fusion leads 22–1
Fugu Ultra 1.1 vs Claude Opus 5
Claude Opus 5 leads 18–4
Fugu Ultra 1.1 vs Hermes MoA
Hermes MoA leads 14–8
Fugu Ultra 1.1 vs GPT-5.6 Sol
GPT-5.6 Sol leads 15–4
Fugu Ultra 1.1 vs Claude Fable 5
Claude Fable 5 leads 16–5
Fugu Ultra 1.1 vs Qwen 3.8
Qwen 3.8 leads 10–6
Fugu Ultra 1.1 vs Grok
Grok leads 14–4
Fugu Ultra 1.1 vs MiniMax M3
MiniMax M3 leads 17–5
Fugu Ultra 1.1 vs Fugu Ultra
Fugu Ultra leads 10–8
Fugu Ultra 1.1 vs Kimi K3
Kimi K3 leads 13–8
Fugu Ultra 1.1 vs GLM-5.2
GLM-5.2 leads 12–11
Fugu Ultra 1.1 vs Fugu Mini
Fugu Mini leads 8–6
Fugu Ultra 1.1 vs Muse Spark 1.2
Fugu Ultra 1.1 leads 11–10
Fugu Ultra 1.1 vs Opus 4.8
Opus 4.8 leads 12–11
Fugu Ultra 1.1 vs Kimi K2.7
Kimi K2.7 leads 6–5
Fugu Ultra 1.1 vs Qwable 5 27B Coder
Fugu Ultra 1.1 leads 10–7
Fugu Ultra 1.1 vs Gemini 3.6 Flash
Tied 11–11
Fugu Ultra 1.1 vs Claude Sonnet 5
Fugu Ultra 1.1 leads 12–10
Fugu Ultra 1.1 vs Qwen 3.7
Fugu Ultra 1.1 leads 12–11
Fugu Ultra 1.1 vs Inkling
Fugu Ultra 1.1 leads 17–5
Fugu Ultra 1.1 vs Grok 4.6
Fugu Ultra 1.1 leads 12–6
Fugu Ultra 1.1 vs Agents-A1
Fugu Ultra 1.1 leads 14–4
Fugu Ultra 1.1 vs Gemma 4 12B · MLX
Fugu Ultra 1.1 leads 15–3
Fugu Ultra 1.1 vs Laguna XS 2.1
Fugu Ultra 1.1 leads 15–2
Fugu Ultra 1.1 vs Qwythos 9B
Fugu Ultra 1.1 leads 16–3
Fugu Ultra 1.1 vs LongCat-2.0
LongCat-2.0 leads 2–1
Fugu Ultra 1.1 vs Hy3
Fugu Ultra 1.1 leads 4–1
Fugu Ultra 1.1 vs Gemma-4 12B Coder
Gemma-4 12B Coder leads 1–0
Fugu Ultra 1.1 vs DeepSeek V4 Flash
24 shared tasks · unscored
Fugu Ultra 1.1 vs DeepSeek V4 Pro
24 shared tasks · unscored
Fugu Ultra 1.1 vs Kimi K2.7 · Fast
24 shared tasks · unscored
Fugu Ultra 1.1 vs Kimi K2.7 · No-Think
24 shared tasks · unscored
Fugu Ultra 1.1 vs Kimi K2.7 · Quality
24 shared tasks · unscored
Fugu Ultra 1.1 vs Ornith 1.0
19 shared tasks · unscored
Fugu Ultra 1.1 vs Claude Mythos 5
Reference-only
Fugu Ultra 1.1 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Fugu Ultra 1.1 vs Fusion Fugu Ultra 1.1 vs Claude Opus 5 Fugu Ultra 1.1 vs Hermes MoA Fugu Ultra 1.1 vs GPT-5.6 Sol Fugu Ultra 1.1 vs Claude Fable 5 Fugu Ultra 1.1 vs Qwen 3.8 Fugu Ultra 1.1 vs Grok Fugu Ultra 1.1 vs MiniMax M3 Fugu Ultra 1.1 vs Fugu Ultra Fugu Ultra 1.1 vs Kimi K3 Fugu Ultra 1.1 vs GLM-5.2 Fugu Ultra 1.1 vs Fugu Mini Fugu Ultra 1.1 vs Muse Spark 1.2 Fugu Ultra 1.1 vs Opus 4.8 Fugu Ultra 1.1 vs Kimi K2.7 Fugu Ultra 1.1 vs Qwable 5 27B Coder Fugu Ultra 1.1 vs Gemini 3.6 Flash Fugu Ultra 1.1 vs Claude Sonnet 5 Fugu Ultra 1.1 vs Qwen 3.7 Fugu Ultra 1.1 vs Inkling Fugu Ultra 1.1 vs Grok 4.6 Fugu Ultra 1.1 vs Agents-A1 Fugu Ultra 1.1 vs Gemma 4 12B · MLX Fugu Ultra 1.1 vs Laguna XS 2.1 Fugu Ultra 1.1 vs Qwythos 9B Fugu Ultra 1.1 vs LongCat-2.0 Fugu Ultra 1.1 vs Hy3 Fugu Ultra 1.1 vs Gemma-4 12B Coder

Read more on agentos.guide:

Fugu Ultra 1.1 — frequently asked

What is Fugu Ultra 1.1?

Fugu Ultra 1.1 is Sakana AI's AI model — Sakana's multi-agent orchestrator, v1.1 — routes experts per request. It has a 1M tokens context window and was released 2026-07.

How good is Fugu Ultra 1.1 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 6.94/10 across 23 scored tasks, with 0 gold, 1 silver and 2 bronze medals.

How much does Fugu Ultra 1.1 cost?

API · orchestration billed. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2

Where can I see Fugu Ultra 1.1 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly