
Real head-to-head · same prompt, one shot
Fugu Ultra vs Gemini 3.6 Flash
Sakana's multi-agent answer to Fusion — frontier ensemble without single-vendor risk. vs Google's launch-day Flash — faster, cheaper, fewer tokens.
Head-to-head verdict: Fugu Ultra wins 28–14.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Fugu Ultra and Gemini 3.6 Flash, side by side, on 42 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Fugu Ultra · Dispatched from Agent OS as the panel-ensemble alternative to OpenRouter Fusion. Bench scored by Claude judge against the same 42 prompts as every other model.
Gemini 3.6 Flash · Benched via the native Gemini API on launch day. Game tasks use the skill-infused threejs-game-director prompt (same as the rest of the field) and are judged on a real mid-play frame by the same Opus vision judge.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Fugu Ultra
Gemini 3.6 Flash
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Page
Page
Sim
Sim
Sim
Where Fugu Ultra beat Gemini 3.6 Flash
The tasks where I gave Fugu Ultra a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Waves
Visual
Fugu Ultra 8.5
·
Gemini 3.6 Flash 3.0
(+5.5)
What I saw: Ultra v2 — Gerstner ocean waves. Smoke-test PASS (3.6% pixel diff).
Blackhole
Sim
Fugu Ultra 8.5
·
Gemini 3.6 Flash 3.5
(+5.0)
What I saw: Ultra v2 — gravitational-lensing black hole. Smoke-test PASS (2.2% pixel diff).
Pathtracer
Sim
Fugu Ultra 8.5
·
Gemini 3.6 Flash 3.5
(+5.0)
What I saw: Ultra v2 — WebGL path tracer with sample accumulation. Smoke-test PASS (4.1% pixel diff).
Particleforge
Sim
Fugu Ultra 8.0
·
Gemini 3.6 Flash 3.5
(+4.5)
What I saw: Ultra v2 (gap-fill) — mouse-gravity particle sculptor. Smoke-test PASS (2.6% diff).
Voxel
Visual
Fugu Ultra 9.0
·
Gemini 3.6 Flash 4.5
(+4.5)
· winner · voxel runner
What I saw: Ultra v2 — Temple-Run voxel runner. Smoke-test PASS with 32.7% pixel diff — the single most reactive build. Replaces the earlier truncated voxel-fugu that was deleted.
Where Gemini 3.6 Flash beat Fugu Ultra
The tasks where I gave Gemini 3.6 Flash a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Dogfight
Game
Gemini 3.6 Flash 8.4
·
Fugu Ultra 7.0
(+1.4)
· polished 3D dogfight
What I saw: Gorgeous low-poly 3D scene with a detailed player jet, glowing engines, volumetric clouds, terrain, and a full sci-fi HUD (hull, boost, radar with hostile blips, target-lock reticle). Enemies are present on radar and in-world but combat isn't clearly shown mid-fight in the shot, …
Outrun
Game
Gemini 3.6 Flash 8.4
·
Fugu Ultra 7.0
(+1.4)
· synthwave combat runner
What I saw: Gorgeous synthwave scene with pseudo-3D grid road, glowing sun, neon palms, polished HUD and an on-road enemy plus reticle showing real combat mechanics; loses a touch because the road perspective looks partly flat/wide rather than a tight curving pseudo-3D outrun feel.
Dragonrealm
Game
Gemini 3.6 Flash 8.3
·
Fugu Ultra 7.0
(+1.3)
What I saw: Strong atmospheric frozen world with snowy terrain, pine forest, night sky, a viewable held sword, a shrine/altar with particle effects, and a visible humanoid enemy plus full HUD (vitality/stamina/compass/kills). Polished and clearly shippable, but the sword FP model looks a bit…
Fluid
Sim
Gemini 3.6 Flash 7.8
·
Fugu Ultra 6.5
(+1.3)
What I saw: Strong particle density (75k) with glowing additive blending, polished glassmorphism UI, presets and shockwave controls all render cleanly; but the center is blown out to solid white and it reads more as a bright particle cloud than an elegant swirling fluid, keeping it just shor…
Webos
Page
Gemini 3.6 Flash 8.0
·
Fugu Ultra 7.0
(+1.0)
What I saw: Strong, polished glassmorphic desktop with a 3D WebGL wireframe background, working top bar, dock (Terminal/Paint/Notes/Settings), and a functional-looking terminal with prompt; the visible Paint window appears empty (only toolbar/slider showing, no canvas content) which slightly…
Strengths & weaknesses I logged
Fugu Ultra
Strengths
- SWE Bench Pro 73.7 · GPQA-D 95.5 · MRCRv2 93.6 — Sakana's published frontier-tier benchmark scores
- Vendor-agnostic ensemble — opt out of specific providers for compliance / export-control
- OpenAI-compatible API at api.sakana.ai — drop-in for existing tooling
Trade-offs
- Panel orchestration adds latency — even a 'pong' burns ~2k orchestration tokens
- Newer than Fusion; less community calibration on long-tail prompts
Gemini 3.6 Flash
Strengths
- Fast one-shot builds — full skill-spec 3D games in ~60-120s of generation
- Cheapest frontier-tier entry on the bench at $1.50/M input
- 17% fewer output tokens than 3.5 Flash on the same workflows (Google's launch claim)
Trade-offs
- Benched on launch day — partial run until the full 50-task batch completes
- Flash tier, not a flagship — up against Pro/flagship-class models on this board
Pricing & context — the spec sheet
| Spec | Fugu Ultra | Gemini 3.6 Flash |
|---|---|---|
| Vendor | Sakana AI | |
| Context window | 272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates. | 1,000,000-token context window |
| Price | $5 / 1M input · $30 / 1M output (Fugu Ultra) | $1.50 / M input |
| Pricing detail | Sakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M output for Fugu Ultra. Direct competitor to OpenRouter Fusion's panel approach. | Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day. |
| Release | 2026-06-15 | 2026-07 |
| Bench coverage | 42/42 scored · avg 7.94/10 | 50/50 scored · avg 7.08/10 |
The verdict — which should you pick?
Across 42 scored shared tasks, Fugu Ultra averaged 7.94/10, beating Gemini 3.6 Flash's 6.92/10 by 1.02 points. Pick Fugu Ultra when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Fugu Ultra and Gemini 3.6 Flash both into the Agent Operating System and dispatch each from the kanban by task type — teams that want fusion-class quality but need a different vendor risk profile → Fugu Ultra, high-volume agentic work where token cost dominates → Gemini 3.6 Flash. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Fugu Ultra vs Gemini 3.6 Flash
Which is better, Fugu Ultra or Gemini 3.6 Flash?
On Goldie Bench, Fugu Ultra averages 7.94/10 across the shared tasks, with 5 gold, 2 silver, 3 bronze overall. Gemini 3.6 Flash averages 6.92/10, with 2 gold, 3 silver, 2 bronze. Fugu Ultra wins the head-to-head 28–14.
How much does Fugu Ultra cost vs Gemini 3.6 Flash?
Fugu Ultra: Sakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M output for Fugu Ultra. Direct competitor to OpenRouter Fusion's panel approach. Gemini 3.6 Flash: Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.
What's the context window for Fugu Ultra vs Gemini 3.6 Flash?
Fugu Ultra has a 272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates. context window. Gemini 3.6 Flash has a 1,000,000-token context window context window.
When should I pick Fugu Ultra over Gemini 3.6 Flash?
Pick Fugu Ultra for: Teams that want Fusion-class quality but need a different vendor risk profile; Operators avoiding export-controlled providers (Sakana emphasises this in their pitch); Deep-research workflows where ensemble verdicts beat single-model answers. The trade-off is the weaknesses we logged on the bench: Panel orchestration adds latency — even a 'pong' burns ~2k orchestration tokens; Newer than Fusion; less community calibration on long-tail prompts.
When should I pick Gemini 3.6 Flash over Fugu Ultra?
Pick Gemini 3.6 Flash for: High-volume agentic work where token cost dominates; Fast prototype builds you iterate on rather than one-shot masterpieces; Routing the everyday 90% while a flagship handles the hard 10%. The trade-off is the weaknesses we logged on the bench: Benched on launch day — partial run until the full 50-task batch completes; Flash tier, not a flagship — up against Pro/flagship-class models on this board.
How does Goldie Bench score Fugu Ultra vs Gemini 3.6 Flash?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Fugu Ultra vs Fusion Gemini 3.6 Flash vs Fusion Fugu Ultra vs Hermes MoA Gemini 3.6 Flash vs Hermes MoA Fugu Ultra vs GPT-5.6 Sol Gemini 3.6 Flash vs GPT-5.6 Sol Fugu Ultra vs Claude Fable 5 Gemini 3.6 Flash vs Claude Fable 5Full model pages: Fugu Ultra · Gemini 3.6 Flash · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































