Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Real head-to-head · same prompt, one shot

Fugu Ultra vs Gemini 3.6 Flash

Sakana's multi-agent answer to Fusion — frontier ensemble without single-vendor risk. vs Google's launch-day Flash — faster, cheaper, fewer tokens.

Head-to-head verdict: Fugu Ultra wins 28–14.

Fugu Ultra · context272K tokens (free) · larger via paid tier
Gemini 3.6 Flash · context1M tokens
Fugu Ultra · price$5 / 1M input · $30 / 1M output (Fugu Ultra)
Gemini 3.6 Flash · price$1.50 / M input
Fugu Ultra · vendorSakana AI
Gemini 3.6 Flash · vendorGoogle

What I tested — same prompt, two models

I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Fugu Ultra and Gemini 3.6 Flash, side by side, on 42 shared tasks inside the Agent Operating System.

Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.

Fugu Ultra · Dispatched from Agent OS as the panel-ensemble alternative to OpenRouter Fusion. Bench scored by Claude judge against the same 42 prompts as every other model.

Gemini 3.6 Flash · Benched via the native Gemini API on launch day. Game tasks use the skill-infused threejs-game-director prompt (same as the rest of the field) and are judged on a real mid-play frame by the same Opus vision judge.

Side-by-side on 50 shared tasks

Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).

Task ↓
Fugu Ultra
Gemini 3.6 Flash
Game
Fugu Ultra on Arcade
Gemini 3.6 Flash on Arcade
Game
Fugu Ultra on Crypt
Gemini 3.6 Flash on Crypt
Game
Fugu Ultra on Dogfight
Gemini 3.6 Flash on Dogfight
Game
Fugu Ultra on Doom
Gemini 3.6 Flash on Doom
🥉Fugu Ultra on Dragonflight
Gemini 3.6 Flash on Dragonflight
Fugu Ultra on Dragonrealm
Gemini 3.6 Flash on Dragonrealm
Game
🥇Fugu Ultra on Game
Gemini 3.6 Flash on Game
Fugu Ultra on Neonblaster
Gemini 3.6 Flash on Neonblaster
Game
Fugu Ultra on Neoncity
Gemini 3.6 Flash on Neoncity
Game
Fugu Ultra on Neonracer
Gemini 3.6 Flash on Neonracer
🥈Fugu Ultra on Nordiccrypt
Gemini 3.6 Flash on Nordiccrypt
Game
Fugu Ultra on Outrun
Gemini 3.6 Flash on Outrun
Game
Fugu Ultra on Pool
Gemini 3.6 Flash on Pool
Game
Fugu Ultra on Racing
Gemini 3.6 Flash on Racing
Game
🥈Fugu Ultra on Raycaster
Gemini 3.6 Flash on Raycaster
Game
Fugu Ultra on Rpg
Gemini 3.6 Flash on Rpg
Game
Fugu Ultra on Skyrim
Gemini 3.6 Flash on Skyrim
🥉Fugu Ultra on Twilightvale
Gemini 3.6 Flash on Twilightvale
Game
Fugu Ultra on Voxelcraft
Gemini 3.6 Flash on Voxelcraft
Page
🥇Fugu Ultra on Landing
Gemini 3.6 Flash on Landing
Page
Fugu Ultra on Webos
Gemini 3.6 Flash on Webos
Sim
Fugu Ultra on Blackhole
Gemini 3.6 Flash on Blackhole
Sim
Fugu Ultra on Boids
Gemini 3.6 Flash on Boids
Sim
Fugu Ultra on Cloth
Gemini 3.6 Flash on Cloth

Where Fugu Ultra beat Gemini 3.6 Flash

The tasks where I gave Fugu Ultra a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Waves Visual
Fugu Ultra 8.5 · Gemini 3.6 Flash 3.0 (+5.5)

What I saw: Ultra v2 — Gerstner ocean waves. Smoke-test PASS (3.6% pixel diff).

Blackhole Sim
Fugu Ultra 8.5 · Gemini 3.6 Flash 3.5 (+5.0)

What I saw: Ultra v2 — gravitational-lensing black hole. Smoke-test PASS (2.2% pixel diff).

Fugu Ultra 8.5 · Gemini 3.6 Flash 3.5 (+5.0)

What I saw: Ultra v2 — WebGL path tracer with sample accumulation. Smoke-test PASS (4.1% pixel diff).

Fugu Ultra 8.0 · Gemini 3.6 Flash 3.5 (+4.5)

What I saw: Ultra v2 (gap-fill) — mouse-gravity particle sculptor. Smoke-test PASS (2.6% diff).

Voxel Visual
Fugu Ultra 9.0 · Gemini 3.6 Flash 4.5 (+4.5) · winner · voxel runner

What I saw: Ultra v2 — Temple-Run voxel runner. Smoke-test PASS with 32.7% pixel diff — the single most reactive build. Replaces the earlier truncated voxel-fugu that was deleted.

Where Gemini 3.6 Flash beat Fugu Ultra

The tasks where I gave Gemini 3.6 Flash a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Dogfight Game
Gemini 3.6 Flash 8.4 · Fugu Ultra 7.0 (+1.4) · polished 3D dogfight

What I saw: Gorgeous low-poly 3D scene with a detailed player jet, glowing engines, volumetric clouds, terrain, and a full sci-fi HUD (hull, boost, radar with hostile blips, target-lock reticle). Enemies are present on radar and in-world but combat isn't clearly shown mid-fight in the shot, …

Outrun Game
Gemini 3.6 Flash 8.4 · Fugu Ultra 7.0 (+1.4) · synthwave combat runner

What I saw: Gorgeous synthwave scene with pseudo-3D grid road, glowing sun, neon palms, polished HUD and an on-road enemy plus reticle showing real combat mechanics; loses a touch because the road perspective looks partly flat/wide rather than a tight curving pseudo-3D outrun feel.

Gemini 3.6 Flash 8.3 · Fugu Ultra 7.0 (+1.3)

What I saw: Strong atmospheric frozen world with snowy terrain, pine forest, night sky, a viewable held sword, a shrine/altar with particle effects, and a visible humanoid enemy plus full HUD (vitality/stamina/compass/kills). Polished and clearly shippable, but the sword FP model looks a bit…

Fluid Sim
Gemini 3.6 Flash 7.8 · Fugu Ultra 6.5 (+1.3)

What I saw: Strong particle density (75k) with glowing additive blending, polished glassmorphism UI, presets and shockwave controls all render cleanly; but the center is blown out to solid white and it reads more as a bright particle cloud than an elegant swirling fluid, keeping it just shor…

Webos Page
Gemini 3.6 Flash 8.0 · Fugu Ultra 7.0 (+1.0)

What I saw: Strong, polished glassmorphic desktop with a 3D WebGL wireframe background, working top bar, dock (Terminal/Paint/Notes/Settings), and a functional-looking terminal with prompt; the visible Paint window appears empty (only toolbar/slider showing, no canvas content) which slightly…

Strengths & weaknesses I logged

Fugu Ultra

Strengths

  • SWE Bench Pro 73.7 · GPQA-D 95.5 · MRCRv2 93.6 — Sakana's published frontier-tier benchmark scores
  • Vendor-agnostic ensemble — opt out of specific providers for compliance / export-control
  • OpenAI-compatible API at api.sakana.ai — drop-in for existing tooling

Trade-offs

  • Panel orchestration adds latency — even a 'pong' burns ~2k orchestration tokens
  • Newer than Fusion; less community calibration on long-tail prompts

Gemini 3.6 Flash

Strengths

  • Fast one-shot builds — full skill-spec 3D games in ~60-120s of generation
  • Cheapest frontier-tier entry on the bench at $1.50/M input
  • 17% fewer output tokens than 3.5 Flash on the same workflows (Google's launch claim)

Trade-offs

  • Benched on launch day — partial run until the full 50-task batch completes
  • Flash tier, not a flagship — up against Pro/flagship-class models on this board

Pricing & context — the spec sheet

Spec Fugu Ultra Gemini 3.6 Flash
VendorSakana AIGoogle
Context window272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates.1,000,000-token context window
Price$5 / 1M input · $30 / 1M output (Fugu Ultra)$1.50 / M input
Pricing detailSakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M output for Fugu Ultra. Direct competitor to OpenRouter Fusion's panel approach.Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.
Release2026-06-152026-07
Bench coverage42/42 scored · avg 7.94/1050/50 scored · avg 7.08/10

The verdict — which should you pick?

Across 42 scored shared tasks, Fugu Ultra averaged 7.94/10, beating Gemini 3.6 Flash's 6.92/10 by 1.02 points. Pick Fugu Ultra when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.

If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Fugu Ultra and Gemini 3.6 Flash both into the Agent Operating System and dispatch each from the kanban by task type — teams that want fusion-class quality but need a different vendor risk profile → Fugu Ultra, high-volume agentic work where token cost dominates → Gemini 3.6 Flash. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.

FAQ — Fugu Ultra vs Gemini 3.6 Flash

Which is better, Fugu Ultra or Gemini 3.6 Flash?

On Goldie Bench, Fugu Ultra averages 7.94/10 across the shared tasks, with 5 gold, 2 silver, 3 bronze overall. Gemini 3.6 Flash averages 6.92/10, with 2 gold, 3 silver, 2 bronze. Fugu Ultra wins the head-to-head 28–14.

How much does Fugu Ultra cost vs Gemini 3.6 Flash?

Fugu Ultra: Sakana's multi-agent orchestration: a single API call internally dispatches to multiple frontier models and synthesises the answer. Subscription plans run $20-$200/mo (Standard / Pro / Max); PAYG is $5/M input + $30/M output for Fugu Ultra. Direct competitor to OpenRouter Fusion's panel approach. Gemini 3.6 Flash: Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.

What's the context window for Fugu Ultra vs Gemini 3.6 Flash?

Fugu Ultra has a 272,000 tokens with the standard rate. Calls exceeding 272K context are billed at the higher 'long-context' rates. context window. Gemini 3.6 Flash has a 1,000,000-token context window context window.

When should I pick Fugu Ultra over Gemini 3.6 Flash?

Pick Fugu Ultra for: Teams that want Fusion-class quality but need a different vendor risk profile; Operators avoiding export-controlled providers (Sakana emphasises this in their pitch); Deep-research workflows where ensemble verdicts beat single-model answers. The trade-off is the weaknesses we logged on the bench: Panel orchestration adds latency — even a 'pong' burns ~2k orchestration tokens; Newer than Fusion; less community calibration on long-tail prompts.

When should I pick Gemini 3.6 Flash over Fugu Ultra?

Pick Gemini 3.6 Flash for: High-volume agentic work where token cost dominates; Fast prototype builds you iterate on rather than one-shot masterpieces; Routing the everyday 90% while a flagship handles the hard 10%. The trade-off is the weaknesses we logged on the bench: Benched on launch day — partial run until the full 50-task batch completes; Flash tier, not a flagship — up against Pro/flagship-class models on this board.

How does Goldie Bench score Fugu Ultra vs Gemini 3.6 Flash?

Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly