Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Gemini 3.6 Flash's full benchmark breakdown.

Google's launch-day Flash — faster, cheaper, fewer tokens. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-07-21
Scored tasks
50
Reading time
10 min
External sources
1

01 · The headline numbers

GoldieBench average
7.08/10
50 scored one-shot tasks
Board rank
#15
of 19 ranked frontier models
Task medals
2🥇 3🥈 2🥉
outright wins on shared briefs
Context
1,000,000-token context window
$1.50 / M input

02 · Every benchmark, bar by bar

All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Stunning volumetric 3D depth with proper katakana glyphs, glowing head chars and light-streak trails in perspective — well beyond a flat 2D canvas rain. Rich UI with palette themes, orbit/zoom controls and EMP pulse pushes it above the field's best.”

- the judge's verdict, unedited

“Strong spiral structure with pink/teal core glow, background stars, and a clean polished UI with palette presets and clear swirl/orbit/zoom controls; the custom shader with mouse interaction and multiple palettes pushes it to the top of the field.”

- the judge's verdict, unedited

“Gorgeous full-screen GLSL plasma with smooth vibrant flowing color bands, polished glassmorphic header and clean palette switcher with 5 presets, plus ripple/drag distortion logic in the shader. Screenshot is genuinely hypnotic and premium; only nit is ripples/distortion can't be verified statically, but the visual quality tops the field.”

- the judge's verdict, unedited

“Strong 3D scene with a well-rendered deployed canopy, articulated diver with suspension lines, layered clouds, jungle terrain and a slick functional HUD (altitude/descent/compass/distance). Slightly held back by the drone-combat framing feeling tacked-on and no enemies visible in this shot, but visually it competes near the top of the field.”

- the judge's verdict, unedited

“Strong low-poly terrain with convincing biome coloring (green slopes, sandy shores, snow peaks, blue water), tree instances, stars and polished glassmorphism UI with 5 biome presets plus auto-roam/regenerate; falls just short of the top only for a somewhat flat lighting and generic navigation feel.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Gemini 3.6 Flash's lowest scored builds, verdicts unedited:

“The UI (title card, preset buttons) renders cleanly and the shader source is ambitious with Gerstner waves and Fresnel, but the actual ocean mesh is completely absent from the screenshot — only scattered particle dots appear, meaning the core wave simulation failed to render (likely a shader/uniform error like normalMat).”

- the judge's verdict, unedited

“Polished UI with full controls (bounces, roughness, DOF, focal, presets) and a plausible path-tracing shader in source, but the render is essentially broken — a flat blue-gray plane with a triangular noise wedge and no visible Cornell box, spheres, or lighting, indicating a camera/geometry or accumulation failure.”

- the judge's verdict, unedited

“Polished glassmorphism UI with presets and sliders, but the core requirement—visible swirling particles—is completely absent in the screenshot, showing only an empty gradient background. The WebGL particle system failed to render, gutting the entire brief.”

- the judge's verdict, unedited

“Polished glassmorphism UI and an ambitious raymarching lensing shader, but the render is badly broken — a white/blown-out background with noisy artifacts and a black dome instead of a proper event horizon with a lensed accretion disk, plus a 2 FPS counter indicating serious performance failure.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (23 tasks)

Gemini 3.6 Flash7.3
field avg7.1

Others (3 tasks)

Gemini 3.6 Flash7.8
field avg7.5

Pages (3 tasks)

Gemini 3.6 Flash8.1
field avg7.6

Sims (12 tasks)

Gemini 3.6 Flash6.2
field avg6.7

Visuals (9 tasks)

Gemini 3.6 Flash7.0
field avg6.8

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Gemini 3.6 Flash perform on GoldieBench?

Gemini 3.6 Flash averages 7.08/10 across 50 scored one-shot build tasks, ranking #15 of 19 ranked frontier models, with 2 task golds, 3 silvers and 2 bronzes. Every score is a real judged build you can open and play on this site.

02What is Gemini 3.6 Flash best at?

Its strongest scored build is Matrix at 8.7/10. By category it averages Game 7.3, Other 7.8, Page 8.1, Sim 6.2, Visual 7.0 on the bench.

03How much does Gemini 3.6 Flash cost?

Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.

Source ledger

  1. 01Official vendor site: ai.google.devai.google.dev

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly