
Real head-to-head · same prompt, one shot
Gemini 3.6 Flash vs Inkling
Google's launch-day Flash — faster, cheaper, fewer tokens. vs A 975B open-weights frontier model — yours to own and run.
Head-to-head verdict: Gemini 3.6 Flash wins 37–7 with 6 ties.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Gemini 3.6 Flash and Inkling, side by side, on 50 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Gemini 3.6 Flash · Benched via the native Gemini API on launch day. Game tasks use the skill-infused threejs-game-director prompt (same as the rest of the field) and are judged on a real mid-play frame by the same Opus vision judge.
Inkling · Benched on GoldieBench one-shot through Tinker's OpenAI-compatible endpoint at medium reasoning effort, then headless-playtested on the same rubric as the whole field. In the Agent OS it's wired into the opencode tab on your own Tinker key — the Ink Machine.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Gemini 3.6 Flash
Inkling
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Other
Where Gemini 3.6 Flash beat Inkling
The tasks where I gave Gemini 3.6 Flash a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Plasma
Visual
Gemini 3.6 Flash 8.6
·
Inkling 2.3
(+6.3)
· WebGL plasma polish
What I saw: Gorgeous full-screen GLSL plasma with smooth vibrant flowing color bands, polished glassmorphic header and clean palette switcher with 5 presets, plus ripple/drag distortion logic in the shader. Screenshot is genuinely hypnotic and premium; only nit is ripples/distortion can't be…
Flightsim
Game
Gemini 3.6 Flash 8.1
·
Inkling 3.5
(+4.6)
What I saw: Strong, cohesive flightsim: clean low-poly terrain with mountains/trees, full HUD (airspeed, altitude, VS, heading tape, artificial horizon, throttle), a visible aircraft on the runway, plus a combat/landing loop. No enemies visible in this frame and it's a mid-runway static shot…
Crypt
Game
Gemini 3.6 Flash 6.8
·
Inkling 2.5
(+4.3)
What I saw: Polished HUD with vitality/stamina/torch bars, viewmodel sword and torch, and clean brick dungeon rendering, but the screenshot shows no visible enemies or combat and the lighting reads flat/brown rather than moody torch-lit — a competent but empty walking sim so far.
Rpg
Game
Gemini 3.6 Flash 7.4
·
Inkling 4.0
(+3.4)
What I saw: Strong 3D atmospheric arena with visible enemies (glowing-eyed shadow creatures), working combat/kill tracking, minimap, inventory slots and polished HUD — but the task asked for a top-down RPG with sprites, and this delivers a first-person 3D combat scene, so it drifts from the …
Raycaster
Game
Gemini 3.6 Flash 7.8
·
Inkling 4.5
(+3.3)
What I saw: Strong polished 3D maze with a detailed plasma weapon model, clean sci-fi HUD, functional circular minimap, and HP/kill tracking; but the screenshot shows plain untextured walls (no Wolfenstein-style texturing) and no enemies visible in view despite the '0/8 hostiles' combat fram…
Where Inkling beat Gemini 3.6 Flash
The tasks where I gave Inkling a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Waves
Visual
Inkling 7.2
·
Gemini 3.6 Flash 3.0
(+4.2)
What I saw: Renders cleanly with a polished cyan gradient ocean, glowing WAVES title, and floating droplet spheres, but the wave surface reads as a smooth blob rather than crisp animated swells and the low camera angle flattens the effect. Competent and on-brief but lacks the specular sparkl…
Voxel
Visual
Inkling 7.2
·
Gemini 3.6 Flash 4.5
(+2.7)
What I saw: Renders a clean 3D voxel terrain with decent lighting, shadows, and polished UI overlay, but the random per-cube color scattering reads as noisy rather than coherent Minecraft-style biomes (grass/dirt/stone layers), and the drag-rotate control is disconnected from the actual auto…
Pathtracer
Sim
Inkling 5.5
·
Gemini 3.6 Flash 3.5
(+2.0)
What I saw: Genuine Monte Carlo path tracer with cosine-weighted sampling and accumulation, and the neon title/UI is polished — but the render is dominated by a blown-out overexposed emissive sphere that washes over everything, the other spheres are barely visible, and the tone-mapping (sqrt…
Pool
Game
Inkling 6.3
·
Gemini 3.6 Flash 4.5
(+1.8)
What I saw: Renders a clean 3D table with rack, pockets, and a decent title overlay, but the table sits small in the frame with heavy vignette wasting most of the screen, and the physics have flaws (cue-ball-only drag, no game rules/scoring). The rack looks slightly offset and the balls are …
Arcade
Game
Inkling 8.2
·
Gemini 3.6 Flash 6.5
(+1.7)
What I saw: A polished 3D Breakout in Three.js with a gorgeous gradient title, glowing rainbow brick wall, paddle/ball follow, trail dots and live score badge — clearly renders and is on-brief. Held back from top spot by the loose 2D collision math on a 3D perspective view (paddle bounce/wal…
Strengths & weaknesses I logged
Gemini 3.6 Flash
Strengths
- Fast one-shot builds — full skill-spec 3D games in ~60-120s of generation
- Cheapest frontier-tier entry on the bench at $1.50/M input
- 17% fewer output tokens than 3.5 Flash on the same workflows (Google's launch claim)
Trade-offs
- Benched on launch day — partial run until the full 50-task batch completes
- Flash tier, not a flagship — up against Pro/flagship-class models on this board
Inkling
Strengths
- Genuinely open-weights — the full 975B model is public on Hugging Face; run it on your own key, no black box
- Best one-shot builds are 2D / animation / web — a matrix-rain that topped its task (8.4), plus arcade, fractal, aurora and a mini web-OS all judged shippable (7.6–8.2)
- Frontier-class agentic coding for an open model — 77.6% SWE-bench Verified, ahead of Nemotron 3 Ultra
- 1M-token context, native multimodal (text/image/audio), and a controllable thinking-effort dial
Trade-offs
- One-shot 3D games are weak — three.js dungeons/racers render a title screen but no playable scene, like most open models (crypt 2.5)
- Physics and particle sims are hit-or-miss — black-hole, plasma and cloth one-shots often render dark or static (2.3–3.5)
- Not the strongest overall — the closed frontier (Fable 5) still tops the raw benchmarks; Inkling trades peak for ownership
Pricing & context — the spec sheet
| Spec | Gemini 3.6 Flash | Inkling |
|---|---|---|
| Vendor | Thinking Machines | |
| Context window | 1,000,000-token context window | 1,000,000 tokens |
| Price | $1.50 / M input | $0.33 / M |
| Pricing detail | Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day. | Inkling is open-weights — a 975B-parameter (41B active) Mixture-of-Experts model whose full weights are public on Hugging Face. You run it on your own key through Tinker's OpenAI-compatible endpoint (usage-based, ~$0.33/M sampling, 50% off at launch), or via Together / Fireworks / Modal / Databricks / Baseten. Benched here one-shot at medium reasoning effort via Tinker. |
| Release | 2026-07 | 2026-07 |
| Bench coverage | 50/50 scored · avg 7.08/10 | 50/50 scored · avg 6.07/10 |
The verdict — which should you pick?
Across 50 scored shared tasks, Gemini 3.6 Flash averaged 7.08/10, beating Inkling's 6.07/10 by 1.02 points. Pick Gemini 3.6 Flash when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Gemini 3.6 Flash and Inkling both into the Agent Operating System and dispatch each from the kanban by task type — high-volume agentic work where token cost dominates → Gemini 3.6 Flash, owning a frontier model instead of renting one — on your own key, pennies per build → Inkling. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Gemini 3.6 Flash vs Inkling
Which is better, Gemini 3.6 Flash or Inkling?
On Goldie Bench, Gemini 3.6 Flash averages 7.08/10 across the shared tasks, with 2 gold, 3 silver, 2 bronze overall. Inkling averages 6.07/10, with 0 gold, 0 silver, 0 bronze. Gemini 3.6 Flash wins the head-to-head 37–7.
How much does Gemini 3.6 Flash cost vs Inkling?
Gemini 3.6 Flash: Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day. Inkling: Inkling is open-weights — a 975B-parameter (41B active) Mixture-of-Experts model whose full weights are public on Hugging Face. You run it on your own key through Tinker's OpenAI-compatible endpoint (usage-based, ~$0.33/M sampling, 50% off at launch), or via Together / Fireworks / Modal / Databricks / Baseten. Benched here one-shot at medium reasoning effort via Tinker.
What's the context window for Gemini 3.6 Flash vs Inkling?
Gemini 3.6 Flash has a 1,000,000-token context window context window. Inkling has a 1,000,000 tokens context window.
When should I pick Gemini 3.6 Flash over Inkling?
Pick Gemini 3.6 Flash for: High-volume agentic work where token cost dominates; Fast prototype builds you iterate on rather than one-shot masterpieces; Routing the everyday 90% while a flagship handles the hard 10%. The trade-off is the weaknesses we logged on the bench: Benched on launch day — partial run until the full 50-task batch completes; Flash tier, not a flagship — up against Pro/flagship-class models on this board.
When should I pick Inkling over Gemini 3.6 Flash?
Pick Inkling for: Owning a frontier model instead of renting one — on your own key, pennies per build; Generative visuals, data-viz and single-file web builds you want one-shot; A customizable open base you can fine-tune on Tinker for your own domain. The trade-off is the weaknesses we logged on the bench: One-shot 3D games are weak — three.js dungeons/racers render a title screen but no playable scene, like most open models (crypt 2.5); Physics and particle sims are hit-or-miss — black-hole, plasma and cloth one-shots often render dark or static (2.3–3.5); Not the strongest overall — the closed frontier (Fable 5) still tops the raw benchmarks; Inkling trades peak for ownership.
How does Goldie Bench score Gemini 3.6 Flash vs Inkling?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Gemini 3.6 Flash vs Fusion Inkling vs Fusion Gemini 3.6 Flash vs Hermes MoA Inkling vs Hermes MoA Gemini 3.6 Flash vs GPT-5.6 Sol Inkling vs GPT-5.6 Sol Gemini 3.6 Flash vs Claude Fable 5 Inkling vs Claude Fable 5Full model pages: Gemini 3.6 Flash · Inkling · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































