Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Real head-to-head · same prompt, one shot

Qwen 3.8 vs Gemini 3.6 Flash

Alibaba's 2.4T flagship — benched through Qoder. vs Google's launch-day Flash — faster, cheaper, fewer tokens.

Head-to-head verdict: Qwen 3.8 wins 35–9 with 1 tie.

Qwen 3.8 · contextQoder-hosted
Gemini 3.6 Flash · context1M tokens
Qwen 3.8 · priceQoder plan
Gemini 3.6 Flash · price$1.50 / M input
Qwen 3.8 · vendorAlibaba
Gemini 3.6 Flash · vendorGoogle

What I tested — same prompt, two models

I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Qwen 3.8 and Gemini 3.6 Flash, side by side, on 45 shared tasks inside the Agent Operating System.

Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.

Qwen 3.8 · Benched on GoldieBench via the Qoder CLI (`qoder-qwen`, model Qwen3.8-Max-Preview) — the only door while it has no public API. Non-game tasks are one-shot like the rest of the field. GAME tasks run in Qoder's real AGENT mode: skill-infused build, then up to 2 QA fix rounds where a vision judge + live console errors are fed back and Qwen 3.8 edits its own file (it took crypt from a black-screen 3.0 to a torch-lit 7.8). Scored on a mid-play frame by the same Opus judge as everyone else.

Gemini 3.6 Flash · Benched via the native Gemini API on launch day. Game tasks use the skill-infused threejs-game-director prompt (same as the rest of the field) and are judged on a real mid-play frame by the same Opus vision judge.

Side-by-side on 50 shared tasks

Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).

Task ↓
Qwen 3.8
Gemini 3.6 Flash
Game
Qwen 3.8 on Arcade
Gemini 3.6 Flash on Arcade
Game
Qwen 3.8 on Crypt
Gemini 3.6 Flash on Crypt
Game
Qwen 3.8 on Dogfight
Gemini 3.6 Flash on Dogfight
Game
Qwen 3.8 on Doom
Gemini 3.6 Flash on Doom
Qwen 3.8 on Dragonflight
Gemini 3.6 Flash on Dragonflight
Qwen 3.8 on Dragonrealm
Gemini 3.6 Flash on Dragonrealm
Game
🥈Qwen 3.8 on Flightsim
Gemini 3.6 Flash on Flightsim
Game
Qwen 3.8 on Game
Gemini 3.6 Flash on Game
Game
🥈Qwen 3.8 on Gtadrive
Gemini 3.6 Flash on Gtadrive
Game
🥉Qwen 3.8 on Gtafoot
Gemini 3.6 Flash on Gtafoot
🥈Qwen 3.8 on Neonblaster
Gemini 3.6 Flash on Neonblaster
Game
Qwen 3.8 on Neonracer
Gemini 3.6 Flash on Neonracer
Qwen 3.8 on Nordiccrypt
Gemini 3.6 Flash on Nordiccrypt
Game
🥇Qwen 3.8 on Outrun
Gemini 3.6 Flash on Outrun
Game
🥉Qwen 3.8 on Racing
Gemini 3.6 Flash on Racing
Game
Qwen 3.8 on Rpg
Gemini 3.6 Flash on Rpg
Game
🥉Qwen 3.8 on Skyrim
Gemini 3.6 Flash on Skyrim
Qwen 3.8 on Twilightvale
Gemini 3.6 Flash on Twilightvale
Other
🥈Qwen 3.8 on Matrixrain
Gemini 3.6 Flash on Matrixrain
🥇Qwen 3.8 on Mlx Speedtest
🥉Gemini 3.6 Flash on Mlx Speedtest
Other
🥉Qwen 3.8 on Neonsnake
🥉Gemini 3.6 Flash on Neonsnake
Page
🥇Qwen 3.8 on Aipbpromo
Gemini 3.6 Flash on Aipbpromo
Page
Qwen 3.8 on Landing
Gemini 3.6 Flash on Landing
Page
Qwen 3.8 on Webos
Gemini 3.6 Flash on Webos

Where Qwen 3.8 beat Gemini 3.6 Flash

The tasks where I gave Qwen 3.8 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Qwen 3.8 8.6 · Gemini 3.6 Flash 3.5 (+5.1) · Physically-correct materials shine

What I saw: Strong: genuine WebGL path tracer with convincing metal/dielectric/lambert spheres, reflections, refraction and soft shadows on a checker floor, plus a polished amber HUD with live SPP/bounce controls. Weak: Monte Carlo noise is still quite grainy and the sky washes to near-flat …

Blackhole Sim
Qwen 3.8 8.4 · Gemini 3.6 Flash 3.5 (+4.9) · Interstellar-style lensing

What I saw: Strong Interstellar-style render: bright photon ring hugging the black silhouette with the accretion disk arcing over the top and secondary lower image reads convincingly as lensing. Slightly let down by the scattered clumpy orbiting particles that look noisy against the clean di…

Qwen 3.8 8.4 · Gemini 3.6 Flash 3.5 (+4.9) · Gorgeous galaxy disc

What I saw: The rendered galaxy disc is genuinely beautiful — dense purple-to-orange particle field with a glowing molten core, gravity reticle, and rich telemetry/formation HUD that nails the sculpting brief. Held just below top tier by HUD overlap on the left (telemetry/controls panels col…

Racing Game
Qwen 3.8 8.6 · Gemini 3.6 Flash 4.2 (+4.4) · Combat Racer Polish

What I saw: Strong render: sleek third-person hovercraft with a banked track, spike-mine hostiles clustered ahead, tire stacks/pillar obstacles, and a gorgeously cohesive sunset HUD with minimap, crosshair, and combat stats. Beats generic walking sims by delivering visible enemies + combat f…

Cloth Sim
Qwen 3.8 8.7 · Gemini 3.6 Flash 4.5 (+4.2) · Elegant draped cloth

What I saw: Gorgeous, dramatic drape with visible weave texture, corner-pinned peaks and the teal sphere reading clearly through the fabric — cinematic lighting and vignette elevate it above the field; only minor nit is the cloth resolution/collision looks slightly coarse at the sphere contact.

Where Gemini 3.6 Flash beat Qwen 3.8

The tasks where I gave Gemini 3.6 Flash a higher 0–10 score on the same prompt — with the actual commentary from my source guides.

Game Game
Gemini 3.6 Flash 8.2 · Qwen 3.8 6.5 (+1.7)

What I saw: Strong cyberpunk 3D combat game with polished HUD, working score/kill tracking (250 pts, 1 kill), reduced HP showing active combat, and a well-modeled 3D robot enemy with atmospheric neon lighting; loses points because only one enemy is visible and the scene reads slightly static…

Dogfight Game
Gemini 3.6 Flash 8.4 · Qwen 3.8 6.8 (+1.6) · polished 3D dogfight

What I saw: Gorgeous low-poly 3D scene with a detailed player jet, glowing engines, volumetric clouds, terrain, and a full sci-fi HUD (hull, boost, radar with hostile blips, target-lock reticle). Enemies are present on radar and in-world but combat isn't clearly shown mid-fight in the shot, …

Waves Visual
Gemini 3.6 Flash 3.0 · Qwen 3.8 2.0 (+1.0)

What I saw: The UI (title card, preset buttons) renders cleanly and the shader source is ambitious with Gerstner waves and Fresnel, but the actual ocean mesh is completely absent from the screenshot — only scattered particle dots appear, meaning the core wave simulation failed to render (lik…

Gemini 3.6 Flash 7.3 · Qwen 3.8 6.4 (+0.9)

What I saw: Strong atmospheric first-person crypt with polished HUD, textured brick walls, torch lighting, and a nicely modeled viewmodel axe; but the screenshot shows no visible enemies/combat and the framing is dominated by a flat wall, so it reads as a walking sim rather than proving the …

Terrain Visual
Gemini 3.6 Flash 8.4 · Qwen 3.8 7.8 (+0.6) · biome terrain explorer

What I saw: Strong low-poly terrain with convincing biome coloring (green slopes, sandy shores, snow peaks, blue water), tree instances, stars and polished glassmorphism UI with 5 biome presets plus auto-roam/regenerate; falls just short of the top only for a somewhat flat lighting and gener…

Strengths & weaknesses I logged

Qwen 3.8

Strengths

  • Skyrim-style open worlds — the Dragon Realm build rendered a lit snowfield, first-person sword and working roaming enemies (7.8)
  • Held up across game genres early — voxel sandbox and Doom raycaster both came out shippable
  • Runs as a real agentic coder inside Qoder (writes + iterates on files), not just a chat model

Trade-offs

  • Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift
  • Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can

Gemini 3.6 Flash

Strengths

  • Fast one-shot builds — full skill-spec 3D games in ~60-120s of generation
  • Cheapest frontier-tier entry on the bench at $1.50/M input
  • 17% fewer output tokens than 3.5 Flash on the same workflows (Google's launch claim)

Trade-offs

  • Benched on launch day — partial run until the full 50-task batch completes
  • Flash tier, not a flagship — up against Pro/flagship-class models on this board

Pricing & context — the spec sheet

Spec Qwen 3.8 Gemini 3.6 Flash
VendorAlibabaGoogle
Context windowServed through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet.1,000,000-token context window
PriceQoder plan$1.50 / M input
Pricing detailQwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot.Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.
Release2026-072026-07
Bench coverage45/45 scored · avg 8.10/1050/50 scored · avg 7.08/10

The verdict — which should you pick?

Across 45 scored shared tasks, Qwen 3.8 averaged 8.10/10, beating Gemini 3.6 Flash's 7.08/10 by 1.02 points. Pick Qwen 3.8 when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.

If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Qwen 3.8 and Gemini 3.6 Flash both into the Agent Operating System and dispatch each from the kanban by task type — one-shot 3d game and world prototypes where atmosphere matters → Qwen 3.8, high-volume agentic work where token cost dominates → Gemini 3.6 Flash. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.

FAQ — Qwen 3.8 vs Gemini 3.6 Flash

Which is better, Qwen 3.8 or Gemini 3.6 Flash?

On Goldie Bench, Qwen 3.8 averages 8.10/10 across the shared tasks, with 9 gold, 11 silver, 5 bronze overall. Gemini 3.6 Flash averages 7.08/10, with 2 gold, 3 silver, 2 bronze. Qwen 3.8 wins the head-to-head 35–9.

How much does Qwen 3.8 cost vs Gemini 3.6 Flash?

Qwen 3.8: Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot. Gemini 3.6 Flash: Launched 2026-07-21 alongside Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. Google's pitch: higher intelligence than its predecessors on coding/ML/knowledge tasks while using 17% fewer output tokens, at a new lower price ($1.50 per million input tokens). Benched here via the native Gemini API on launch day.

What's the context window for Qwen 3.8 vs Gemini 3.6 Flash?

Qwen 3.8 has a Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. context window. Gemini 3.6 Flash has a 1,000,000-token context window context window.

When should I pick Qwen 3.8 over Gemini 3.6 Flash?

Pick Qwen 3.8 for: One-shot 3D game and world prototypes where atmosphere matters; Anyone already in the Qoder IDE/CLI wanting a near-frontier model free on the Pro trial; A cheaper stand-in for Fable 5 on creative-visual builds. The trade-off is the weaknesses we logged on the bench: Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift; Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can.

When should I pick Gemini 3.6 Flash over Qwen 3.8?

Pick Gemini 3.6 Flash for: High-volume agentic work where token cost dominates; Fast prototype builds you iterate on rather than one-shot masterpieces; Routing the everyday 90% while a flagship handles the hard 10%. The trade-off is the weaknesses we logged on the bench: Benched on launch day — partial run until the full 50-task batch completes; Flash tier, not a flagship — up against Pro/flagship-class models on this board.

How does Goldie Bench score Qwen 3.8 vs Gemini 3.6 Flash?

Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly