Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Sakana AI

Fugu Ultra 1.1

Sakana's multi-agent orchestrator, v1.1 — routes experts per request.

Context1,000,000-token context window
PricingAPI · orchestration billed
Tasks tested0
Avg scorecurrently unranked
Medals🥇0 🥈0 🥉0
Release2026-07
Official sitesakana.ai ↗
Official vendor source
Fugu Ultra 1.1 is built by Sakana AI — see the vendor's own product page, pricing, and docs at sakana.ai.
Visit sakana.ai →

Reference benchmarks for Fugu Ultra 1.1

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Fugu Ultra 1.1 is honest about what's measured.

SWE-Bench Pro (reported)
73.7
source: /sakana.ai
LiveCodeBench (reported)
93.2
source: /sakana.ai
GPQA Diamond (reported)
95.5
source: /sakana.ai

What is Fugu Ultra 1.1?

Fugu Ultra 1.1 is the Sakana AI frontier model with a 1,000,000-token context window context window, released 2026-07. Tagline: Sakana's multi-agent orchestrator, v1.1 — routes experts per request.. Official source: sakana.ai.

Pricing detail. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.

How I use it inside the Agent OS. Benched on GoldieBench via Sakana's Responses API (fugu-ultra-v1.1, xhigh reasoning). Game tasks use the skill-infused threejs-game-director prompt plus a controls+graphics fix loop with an anti-regression clamp, judged on a real mid-play frame by the same Opus judge as the field.

What I built with Fugu Ultra 1.1

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Fugu Ultra 1.1 shipped on the bench: 0 one-shot demos across 1,000,000-token context window of context. Of those, 0 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Orchestrates 1-3 expert agents per request and synthesises their answers
  • Reported SWE-Bench Pro 73.7 — above Opus 4.8 and GPT-5.5 on Sakana's table
  • OpenAI- and Anthropic-compatible API — drop-in for Codex and Claude Code

Trade-offs

  • Benched as a partial run until the full 50-task batch completes
  • Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are

Best for

  • Hard, high-stakes coding and reasoning where answer quality beats latency
  • Agentic workflows in Codex / Claude Code via the drop-in provider config
  • One-shot builds you want a panel of experts on, not a single model

Every demo by Fugu Ultra 1.1

0 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

every demo, in a grid · click any one to play

Compare Fugu Ultra 1.1 against every other model

Every head-to-head featuring Fugu Ultra 1.1. Verdicts shown for scored pairs.

Fugu Ultra 1.1 vs Fusion
Reference-only
Fugu Ultra 1.1 vs Hermes MoA
Reference-only
Fugu Ultra 1.1 vs GPT-5.6 Sol
Reference-only
Fugu Ultra 1.1 vs Claude Fable 5
Reference-only
Fugu Ultra 1.1 vs Qwen 3.8
Reference-only
Fugu Ultra 1.1 vs Grok
Reference-only
Fugu Ultra 1.1 vs MiniMax M3
Reference-only
Fugu Ultra 1.1 vs Fugu Ultra
Reference-only
Fugu Ultra 1.1 vs Kimi K3
Reference-only
Fugu Ultra 1.1 vs GLM-5.2
Reference-only
Fugu Ultra 1.1 vs Fugu Mini
Reference-only
Fugu Ultra 1.1 vs Opus 4.8
Reference-only
Fugu Ultra 1.1 vs Kimi K2.7
Reference-only
Fugu Ultra 1.1 vs Qwable 5 27B Coder
Reference-only
Fugu Ultra 1.1 vs Gemini 3.6 Flash
Reference-only
Fugu Ultra 1.1 vs Claude Sonnet 5
Reference-only
Fugu Ultra 1.1 vs Qwen 3.7
Reference-only
Fugu Ultra 1.1 vs Inkling
Reference-only
Fugu Ultra 1.1 vs Claude Opus 5
Reference-only
Fugu Ultra 1.1 vs Agents-A1
Reference-only
Fugu Ultra 1.1 vs Gemma 4 12B · MLX
Reference-only
Fugu Ultra 1.1 vs Laguna XS 2.1
Reference-only
Fugu Ultra 1.1 vs Qwythos 9B
Reference-only
Fugu Ultra 1.1 vs LongCat-2.0
Reference-only
Fugu Ultra 1.1 vs Hy3
Reference-only
Fugu Ultra 1.1 vs Gemma-4 12B Coder
Reference-only
Fugu Ultra 1.1 vs Kimi K2.7 · Fast
Reference-only
Fugu Ultra 1.1 vs Kimi K2.7 · No-Think
Reference-only
Fugu Ultra 1.1 vs Kimi K2.7 · Quality
Reference-only
Fugu Ultra 1.1 vs Ornith 1.0
Reference-only
Fugu Ultra 1.1 vs Claude Mythos 5
Reference-only
Fugu Ultra 1.1 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Fugu Ultra 1.1 vs Fusion Fugu Ultra 1.1 vs Hermes MoA Fugu Ultra 1.1 vs GPT-5.6 Sol Fugu Ultra 1.1 vs Claude Fable 5 Fugu Ultra 1.1 vs Qwen 3.8 Fugu Ultra 1.1 vs Grok Fugu Ultra 1.1 vs MiniMax M3 Fugu Ultra 1.1 vs Fugu Ultra Fugu Ultra 1.1 vs Kimi K3 Fugu Ultra 1.1 vs GLM-5.2 Fugu Ultra 1.1 vs Fugu Mini Fugu Ultra 1.1 vs Opus 4.8 Fugu Ultra 1.1 vs Kimi K2.7 Fugu Ultra 1.1 vs Qwable 5 27B Coder Fugu Ultra 1.1 vs Gemini 3.6 Flash Fugu Ultra 1.1 vs Claude Sonnet 5 Fugu Ultra 1.1 vs Qwen 3.7 Fugu Ultra 1.1 vs Inkling Fugu Ultra 1.1 vs Claude Opus 5 Fugu Ultra 1.1 vs Agents-A1 Fugu Ultra 1.1 vs Gemma 4 12B · MLX Fugu Ultra 1.1 vs Laguna XS 2.1 Fugu Ultra 1.1 vs Qwythos 9B Fugu Ultra 1.1 vs LongCat-2.0 Fugu Ultra 1.1 vs Hy3 Fugu Ultra 1.1 vs Gemma-4 12B Coder

Read more on agentos.guide:

Fugu Ultra 1.1 — frequently asked

What is Fugu Ultra 1.1?

Fugu Ultra 1.1 is Sakana AI's AI model — Sakana's multi-agent orchestrator, v1.1 — routes experts per request. It has a 1M tokens context window and was released 2026-07.

How good is Fugu Ultra 1.1 at coding and one-shot builds?

It has 0 live demos on GoldieBench but no curated 0-10 verdicts yet — it is unranked until scored.

How much does Fugu Ultra 1.1 cost?

API · orchestration billed. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2

Where can I see Fugu Ultra 1.1 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly