
Real head-to-head · same prompt, one shot
Qwen 3.8 vs Gemini 3.6 Flash
Alibaba's 2.4T flagship — benched through Qoder. vs Google's launch-day Flash — faster, cheaper, fewer tokens.
Head-to-head verdict: Qwen 3.8 wins 35–9 with 1 tie.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Qwen 3.8 and Gemini 3.6 Flash, side by side, on 45 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Qwen 3.8 · Benched on GoldieBench via the Qoder CLI (`qoder-qwen`, model Qwen3.8-Max-Preview) — the only door while it has no public API. Non-game tasks are one-shot like the rest of the field. GAME tasks run in Qoder's real AGENT mode: skill-infused build, then up to 2 QA fix rounds where a vision judge + live console errors are fed back and Qwen 3.8 edits its own file (it took crypt from a black-screen 3.0 to a torch-lit 7.8). Scored on a mid-play frame by the same Opus judge as everyone else.
Gemini 3.6 Flash · Benched via the native Gemini API on launch day. Game tasks use the skill-infused threejs-game-director prompt (same as the rest of the field) and are judged on a real mid-play frame by the same Opus vision judge.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Qwen 3.8
Gemini 3.6 Flash
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Other
Other
Other
Page
Page
Page
Where Qwen 3.8 beat Gemini 3.6 Flash
The tasks where I gave Qwen 3.8 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Pathtracer
Sim
Qwen 3.8 8.6
·
Gemini 3.6 Flash 3.5
(+5.1)
· Physically-correct materials shine
What I saw: Strong: genuine WebGL path tracer with convincing metal/dielectric/lambert spheres, reflections, refraction and soft shadows on a checker floor, plus a polished amber HUD with live SPP/bounce controls. Weak: Monte Carlo noise is still quite grainy and the sky washes to near-flat …
Blackhole
Sim
Qwen 3.8 8.4
·
Gemini 3.6 Flash 3.5
(+4.9)
· Interstellar-style lensing
What I saw: Strong Interstellar-style render: bright photon ring hugging the black silhouette with the accretion disk arcing over the top and secondary lower image reads convincingly as lensing. Slightly let down by the scattered clumpy orbiting particles that look noisy against the clean di…
Particleforge
Sim
Qwen 3.8 8.4
·
Gemini 3.6 Flash 3.5
(+4.9)
· Gorgeous galaxy disc
What I saw: The rendered galaxy disc is genuinely beautiful — dense purple-to-orange particle field with a glowing molten core, gravity reticle, and rich telemetry/formation HUD that nails the sculpting brief. Held just below top tier by HUD overlap on the left (telemetry/controls panels col…
Racing
Game
Qwen 3.8 8.6
·
Gemini 3.6 Flash 4.2
(+4.4)
· Combat Racer Polish
What I saw: Strong render: sleek third-person hovercraft with a banked track, spike-mine hostiles clustered ahead, tire stacks/pillar obstacles, and a gorgeously cohesive sunset HUD with minimap, crosshair, and combat stats. Beats generic walking sims by delivering visible enemies + combat f…
Cloth
Sim
Qwen 3.8 8.7
·
Gemini 3.6 Flash 4.5
(+4.2)
· Elegant draped cloth
What I saw: Gorgeous, dramatic drape with visible weave texture, corner-pinned peaks and the teal sphere reading clearly through the fabric — cinematic lighting and vignette elevate it above the field; only minor nit is the cloth resolution/collision looks slightly coarse at the sphere contact.
Where Gemini 3.6 Flash beat Qwen 3.8
The tasks where I gave Gemini 3.6 Flash a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Game
Game
Gemini 3.6 Flash 8.2
·
Qwen 3.8 6.5
(+1.7)
What I saw: Strong cyberpunk 3D combat game with polished HUD, working score/kill tracking (250 pts, 1 kill), reduced HP showing active combat, and a well-modeled 3D robot enemy with atmospheric neon lighting; loses points because only one enemy is visible and the scene reads slightly static…
Dogfight
Game
Gemini 3.6 Flash 8.4
·
Qwen 3.8 6.8
(+1.6)
· polished 3D dogfight
What I saw: Gorgeous low-poly 3D scene with a detailed player jet, glowing engines, volumetric clouds, terrain, and a full sci-fi HUD (hull, boost, radar with hostile blips, target-lock reticle). Enemies are present on radar and in-world but combat isn't clearly shown mid-fight in the shot, …
Waves
Visual
Gemini 3.6 Flash 3.0
·
Qwen 3.8 2.0
(+1.0)
What I saw: The UI (title card, preset buttons) renders cleanly and the shader source is ambitious with Gerstner waves and Fresnel, but the actual ocean mesh is completely absent from the screenshot — only scattered particle dots appear, meaning the core wave simulation failed to render (lik…
Nordiccrypt
Game
Gemini 3.6 Flash 7.3
·
Qwen 3.8 6.4
(+0.9)
What I saw: Strong atmospheric first-person crypt with polished HUD, textured brick walls, torch lighting, and a nicely modeled viewmodel axe; but the screenshot shows no visible enemies/combat and the framing is dominated by a flat wall, so it reads as a walking sim rather than proving the …
Terrain
Visual
Gemini 3.6 Flash 8.4
·
Qwen 3.8 7.8
(+0.6)
· biome terrain explorer
What I saw: Strong low-poly terrain with convincing biome coloring (green slopes, sandy shores, snow peaks, blue water), tree instances, stars and polished glassmorphism UI with 5 biome presets plus auto-roam/regenerate; falls just short of the top only for a somewhat flat lighting and gener…
Strengths & weaknesses I logged
Qwen 3.8
Strengths
- Skyrim-style open worlds — the Dragon Realm build rendered a lit snowfield, first-person sword and working roaming enemies (7.8)
- Held up across game genres early — voxel sandbox and Doom raycaster both came out shippable
- Runs as a real agentic coder inside Qoder (writes + iterates on files), not just a chat model
Trade-offs
- Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift
- Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can
Gemini 3.6 Flash
Strengths
- Fast one-shot builds — full skill-spec 3D games in ~60-120s of generation
- Cheapest frontier-tier entry on the bench at $1.50/M input
- 17% fewer output tokens than 3.5 Flash on the same workflows (Google's launch claim)
Trade-offs
- Benched on launch day — partial run until the full 50-task batch completes
- Flash tier, not a flagship — up against Pro/flagship-class models on this board
Pricing & context — the spec sheet
| Spec | Qwen 3.8 | Gemini 3.6 Flash |
|---|---|---|
| Vendor | Alibaba | |
| Context window | Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. | 1,000,000-token context window |
| Price | Qoder plan | $1.50 / M input |
| Pricing detail | Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot. | Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day. |
| Release | 2026-07 | 2026-07 |
| Bench coverage | 45/45 scored · avg 8.10/10 | 50/50 scored · avg 7.08/10 |
The verdict — which should you pick?
Across 45 scored shared tasks, Qwen 3.8 averaged 8.10/10, beating Gemini 3.6 Flash's 7.08/10 by 1.02 points. Pick Qwen 3.8 when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Qwen 3.8 and Gemini 3.6 Flash both into the Agent Operating System and dispatch each from the kanban by task type — one-shot 3d game and world prototypes where atmosphere matters → Qwen 3.8, high-volume agentic work where token cost dominates → Gemini 3.6 Flash. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Qwen 3.8 vs Gemini 3.6 Flash
Which is better, Qwen 3.8 or Gemini 3.6 Flash?
On Goldie Bench, Qwen 3.8 averages 8.10/10 across the shared tasks, with 9 gold, 11 silver, 5 bronze overall. Gemini 3.6 Flash averages 7.08/10, with 2 gold, 3 silver, 2 bronze. Qwen 3.8 wins the head-to-head 35–9.
How much does Qwen 3.8 cost vs Gemini 3.6 Flash?
Qwen 3.8: Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot. Gemini 3.6 Flash: Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.
What's the context window for Qwen 3.8 vs Gemini 3.6 Flash?
Qwen 3.8 has a Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. context window. Gemini 3.6 Flash has a 1,000,000-token context window context window.
When should I pick Qwen 3.8 over Gemini 3.6 Flash?
Pick Qwen 3.8 for: One-shot 3D game and world prototypes where atmosphere matters; Anyone already in the Qoder IDE/CLI wanting a near-frontier model free on the Pro trial; A cheaper stand-in for Fable 5 on creative-visual builds. The trade-off is the weaknesses we logged on the bench: Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift; Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can.
When should I pick Gemini 3.6 Flash over Qwen 3.8?
Pick Gemini 3.6 Flash for: High-volume agentic work where token cost dominates; Fast prototype builds you iterate on rather than one-shot masterpieces; Routing the everyday 90% while a flagship handles the hard 10%. The trade-off is the weaknesses we logged on the bench: Benched on launch day — partial run until the full 50-task batch completes; Flash tier, not a flagship — up against Pro/flagship-class models on this board.
How does Goldie Bench score Qwen 3.8 vs Gemini 3.6 Flash?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Qwen 3.8 vs Fusion Gemini 3.6 Flash vs Fusion Qwen 3.8 vs Hermes MoA Gemini 3.6 Flash vs Hermes MoA Qwen 3.8 vs GPT-5.6 Sol Gemini 3.6 Flash vs GPT-5.6 Sol Qwen 3.8 vs Claude Fable 5 Gemini 3.6 Flash vs Claude Fable 5Full model pages: Qwen 3.8 · Gemini 3.6 Flash · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































