Fugu Ultra 1.1
Sakana's multi-agent orchestrator, v1.1 — routes experts per request.
Reference benchmarks for Fugu Ultra 1.1
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Fugu Ultra 1.1 is honest about what's measured.
What is Fugu Ultra 1.1?
Fugu Ultra 1.1 is the Sakana AI frontier model with a 1,000,000-token context window context window, released 2026-07. Tagline: Sakana's multi-agent orchestrator, v1.1 — routes experts per request.. Official source: sakana.ai.
Pricing detail. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.
How I use it inside the Agent OS. Benched on GoldieBench via Sakana's Responses API (fugu-ultra-v1.1, xhigh reasoning). Game tasks use the skill-infused threejs-game-director prompt plus a controls+graphics fix loop with an anti-regression clamp, judged on a real mid-play frame by the same Opus judge as the field.
What I built with Fugu Ultra 1.1
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Fugu Ultra 1.1 shipped on the bench: 0 one-shot demos across 1,000,000-token context window of context. Of those, 0 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Orchestrates 1-3 expert agents per request and synthesises their answers
- Reported SWE-Bench Pro 73.7 — above Opus 4.8 and GPT-5.5 on Sakana's table
- OpenAI- and Anthropic-compatible API — drop-in for Codex and Claude Code
Trade-offs
- Benched as a partial run until the full 50-task batch completes
- Region-locked: unavailable across the EU/EEA/UK/CH — access depends on where you are
Best for
- Hard, high-stakes coding and reasoning where answer quality beats latency
- Agentic workflows in Codex / Claude Code via the drop-in provider config
- One-shot builds you want a panel of experts on, not a single model
Every demo by Fugu Ultra 1.1
0 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
Compare Fugu Ultra 1.1 against every other model
Every head-to-head featuring Fugu Ultra 1.1. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Fugu Ultra 1.1 vs Fusion Fugu Ultra 1.1 vs Hermes MoA Fugu Ultra 1.1 vs GPT-5.6 Sol Fugu Ultra 1.1 vs Claude Fable 5 Fugu Ultra 1.1 vs Qwen 3.8 Fugu Ultra 1.1 vs Grok Fugu Ultra 1.1 vs MiniMax M3 Fugu Ultra 1.1 vs Fugu Ultra Fugu Ultra 1.1 vs Kimi K3 Fugu Ultra 1.1 vs GLM-5.2 Fugu Ultra 1.1 vs Fugu Mini Fugu Ultra 1.1 vs Opus 4.8 Fugu Ultra 1.1 vs Kimi K2.7 Fugu Ultra 1.1 vs Qwable 5 27B Coder Fugu Ultra 1.1 vs Gemini 3.6 Flash Fugu Ultra 1.1 vs Claude Sonnet 5 Fugu Ultra 1.1 vs Qwen 3.7 Fugu Ultra 1.1 vs Inkling Fugu Ultra 1.1 vs Claude Opus 5 Fugu Ultra 1.1 vs Agents-A1 Fugu Ultra 1.1 vs Gemma 4 12B · MLX Fugu Ultra 1.1 vs Laguna XS 2.1 Fugu Ultra 1.1 vs Qwythos 9B Fugu Ultra 1.1 vs LongCat-2.0 Fugu Ultra 1.1 vs Hy3 Fugu Ultra 1.1 vs Gemma-4 12B CoderRead more on agentos.guide:
Fugu Ultra 1.1 — frequently asked
What is Fugu Ultra 1.1?
Fugu Ultra 1.1 is Sakana AI's AI model — Sakana's multi-agent orchestrator, v1.1 — routes experts per request. It has a 1M tokens context window and was released 2026-07.
How good is Fugu Ultra 1.1 at coding and one-shot builds?
It has 0 live demos on GoldieBench but no curated 0-10 verdicts yet — it is unranked until scored.
How much does Fugu Ultra 1.1 cost?
API · orchestration billed. Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2
Where can I see Fugu Ultra 1.1 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.