Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Fugu Ultra 1.1's full benchmark breakdown.

Sakana's multi-agent orchestrator, v1.1 — routes experts per request. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-08-13
Scored tasks
23
Reading time
7 min
External sources
2

01 · The headline numbers

GoldieBench average
6.94/10
23 scored one-shot tasks
Board rank
#20
of 23 ranked frontier models
Task medals
0🥇 1🥈 2🥉
outright wins on shared briefs
Context
1,000,000-token context window
API · orchestration billed

02 · Every benchmark, bar by bar

All 23 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Strong deployed-chute skydiver over a jungle canopy with a polished HUD (altitude, dist-to-H, score) plus active drone enemies, flare combat, and 'THREAT DOWN' kill feedback — it delivers the full jump/steer/land loop AND working combat, edging past a bland walking sim.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Fugu Ultra 1.1's lowest scored builds, verdicts unedited:

“The HUD (health, charge, minimap frame, controls) renders but the entire 3D scene is black with no visible maze, walls, enemies, or floor — the raycaster world failed to render. HOSTILES 0/0 and empty minimap confirm the core build is broken/non-rendering.”

- the judge's verdict, unedited

“HUD panels (health/boost/kills/crosshair) render nicely but the entire 3D scene is black — no city, road, car, or enemies visible, indicating the Three.js world failed to render. A non-functional walking/driving sim regardless of the ambitious source.”

- the judge's verdict, unedited

“Visually polished 3D pool-hall arena with a giant table, pockets, minimap and combat HUD, but this is an arena shooter reskin ('cue-knight', enemies, HP/boost) — not a physically simulated billiards game as the brief demands, so it misses the core task entirely despite decent rendering.”

- the judge's verdict, unedited

“Strong HUD, first-person weapon/shield and visible enemies (7) with combat scaffolding are present, but the scene reads bright and flat rather than torch-lit crypt — lighting/atmosphere totally missing the Nordic dungeon mood, and floating stretched geometry looks buggy.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (22 tasks)

Fugu Ultra 1.17.0
field avg7.1

Others (0 tasks)

field avg7.5

Pages (0 tasks)

field avg7.6

Sims (0 tasks)

field avg6.9

Visuals (1 tasks)

Fugu Ultra 1.15.5
field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

From the source guides

MeasureResultSource guide
SWE-Bench Pro (reported)73.7/sakana.ai
LiveCodeBench (reported)93.2/sakana.ai
GPQA Diamond (reported)95.5/sakana.ai

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Fugu Ultra 1.1 perform on GoldieBench?

Fugu Ultra 1.1 averages 6.94/10 across 23 scored one-shot build tasks, ranking #20 of 23 ranked frontier models, with 0 task golds, 1 silvers and 2 bronzes. Every score is a real judged build you can open and play on this site.

02What is Fugu Ultra 1.1 best at?

Its strongest scored build is Dragonrealm at 8.6/10. By category it averages Game 7.0, Visual 5.5 on the bench.

03How much does Fugu Ultra 1.1 cost?

Sakana's Fugu Ultra v1.1 — a multi-agent system served as a model: an orchestrator routes each request across one to three expert agents and synthesises the answer. Reported (v1.0): SWE-Bench Pro 73.7, LiveCodeBench 93.2, GPQA Diamond 95.5. NOTE: Sakana geo-blocks the EU/EEA/UK/Switzerland pending GDPR compliance — benched from an allowed region.

Source ledger

  1. 01Official vendor site: sakana.aisakana.ai
  2. 02SWE-Bench Pro (reported)agentos.guide

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly