GoldieBench deep dive
Fugu Ultra 1.1's full benchmark breakdown.
Sakana's multi-agent orchestrator, v1.1 — routes experts per request. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 23 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Strong deployed-chute skydiver over a jungle canopy with a polished HUD (altitude, dist-to-H, score) plus active drone enemies, flare combat, and 'THREAT DOWN' kill feedback — it delivers the full jump/steer/land loop AND working combat, edging past a bland walking sim.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Fugu Ultra 1.1's lowest scored builds, verdicts unedited:
“The HUD (health, charge, minimap frame, controls) renders but the entire 3D scene is black with no visible maze, walls, enemies, or floor — the raycaster world failed to render. HOSTILES 0/0 and empty minimap confirm the core build is broken/non-rendering.”
“HUD panels (health/boost/kills/crosshair) render nicely but the entire 3D scene is black — no city, road, car, or enemies visible, indicating the Three.js world failed to render. A non-functional walking/driving sim regardless of the ambitious source.”
“Visually polished 3D pool-hall arena with a giant table, pockets, minimap and combat HUD, but this is an arena shooter reskin ('cue-knight', enemies, HP/boost) — not a physically simulated billiards game as the brief demands, so it misses the core task entirely despite decent rendering.”
“Strong HUD, first-person weapon/shield and visible enemies (7) with combat scaffolding are present, but the scene reads bright and flat rather than torch-lit crypt — lighting/atmosphere totally missing the Nordic dungeon mood, and floating stretched geometry looks buggy.”
05 · Category breakdown vs the whole field
Games (22 tasks)
Others (0 tasks)
Pages (0 tasks)
Sims (0 tasks)
Visuals (1 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
From the source guides
| Measure | Result | Source guide |
|---|---|---|
| SWE-Bench Pro (reported) | 73.7 | /sakana.ai |
| LiveCodeBench (reported) | 93.2 | /sakana.ai |
| GPQA Diamond (reported) | 95.5 | /sakana.ai |
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Fugu Ultra 1.1 perform on GoldieBench?
Fugu Ultra 1.1 averages 6.94/10 across 23 scored one-shot build tasks, ranking #20 of 23 ranked frontier models, with 0 task golds, 1 silvers and 2 bronzes. Every score is a real judged build you can open and play on this site.
02What is Fugu Ultra 1.1 best at?
Its strongest scored build is Dragonrealm at 8.6/10. By category it averages Game 7.0, Visual 5.5 on the bench.
03How much does Fugu Ultra 1.1 cost?
Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.
Source ledger
- 01Official vendor site: sakana.aisakana.ai
- 02SWE-Bench Pro (reported)agentos.guide
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.