Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Qwen 3.8's full benchmark breakdown.

Alibaba's 2.4T flagship — benched through Qoder. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-07-20
Scored tasks
41
Reading time
9 min
External sources
1

01 · The headline numbers

GoldieBench average
8.22/10
41 scored one-shot tasks
Board rank
#2
of 18 ranked frontier models
Task medals
10🥇 9🥈 5🥉
outright wins on shared briefs
Context
Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet.
Qoder plan

Provisional flag. Part of this model's run hit infrastructure failures that are being re-run; its average is still settling. Scores shown are real judged builds only.

02 · Every benchmark, bar by bar

All 41 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Gorgeous multi-burst scene with dense sparks, long flowing trails, hue-tinted flashes, and a full city skyline with lit windows plus mirrored water reflections — the teal/magenta glow band and shimmer sell the harbor atmosphere. Strong interactivity (click/tap + auto-fire) and varied explosion types; this is a top-tier fireworks build.”

- the judge's verdict, unedited

“Gorgeous, dramatic drape with visible weave texture, corner-pinned peaks and the teal sphere reading clearly through the fabric — cinematic lighting and vignette elevate it above the field; only minor nit is the cloth resolution/collision looks slightly coarse at the sphere contact.”

- the judge's verdict, unedited

“Stunning GPU Mandelbrot with smooth iteration coloring, distance-estimation rim shading and a gorgeous editorial HUD (readouts, palettes, tour) — the deep spiral detail renders beautifully. Strong on both polish and interactivity; only nit is the somewhat washed-out interior gradient, but it clearly competes for the top of the field.”

- the judge's verdict, unedited

“Strong: a genuinely beautiful, dense spiral with a glowing gold core, color-graded arms (gold→pink→blue) and a polished editorial HUD with live telemetry and interaction hints. Weak: the low 20 fps telemetry hints at heavy load, but the render itself is gorgeous and clearly on-brief, edging past the field's best.”

- the judge's verdict, unedited

“Gorgeous rendering — convincing metaball wax with glowing blobs rising in a beautifully detailed brass-and-glass lamp, plus Monoton neon title, live temp readouts, heat slider and theme chips. Polished, on-brief, and clearly at the top of the field; only nitpick is empty left-side negative space.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Qwen 3.8's lowest scored builds, verdicts unedited:

“Polished HUD (Cinzel headers, vigor/breath meters, objective, axe crosshair) and the scene renders a coherent low-poly ruin with columns, arches and torch posts, but the lighting reads flat/dim rather than genuinely 'torch-lit,' the space feels sparse, and crucially no visible enemies appear (0/5 slain) so the combat core is unproven — works but generic against a 9.5 best.”

- the judge's verdict, unedited

“Gorgeous atmospheric desert render with polished HUD, custom crosshair, and a full combat/wave framework in code — but the screenshot shows a chaotic black jumble of geometry (likely the player craft rendering wrong/clipped) instead of a readable ship or visible enemies, and hull already at 010 with 0 kills suggests broken presentation. Strong art direction, but the central subject looks glitched rather than shippable.”

- the judge's verdict, unedited

“Strong neon aesthetic with polished HUD, CRT vignette, and clean paused panel, but the screenshot shows an empty playfield with no visible snake or food, so the actual gameplay isn't demonstrated. Presentation is above-average yet it reads as a shell rather than a proven, running game.”

- the judge's verdict, unedited

“Excellent military HUD polish—compass tape, radar with contacts, hull/thr/boost bars, and a well-modeled aircraft all render cleanly. But the screenshot shows no visible enemies in the sky and zero combat happening (KILLS 00, empty airspace), leaving it feeling more like a polished flight sim than a proven dogfight.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (18 tasks)

Qwen 3.87.9
field avg7.1

Others (3 tasks)

Qwen 3.87.9
field avg7.5

Pages (2 tasks)

Qwen 3.88.4
field avg7.5

Sims (11 tasks)

Qwen 3.88.6
field avg6.7

Visuals (7 tasks)

Qwen 3.88.5
field avg6.8

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Qwen 3.8 perform on GoldieBench?

Qwen 3.8 averages 8.22/10 across 41 scored one-shot build tasks, ranking #2 of 18 ranked frontier models, with 10 task golds, 9 silvers and 5 bronzes. Every score is a real judged build you can open and play on this site.

02What is Qwen 3.8 best at?

Its strongest scored build is Fireworks at 9.0/10. By category it averages Game 7.9, Other 7.9, Page 8.4, Sim 8.6, Visual 8.5 on the bench.

03How much does Qwen 3.8 cost?

Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot.

Source ledger

  1. 01Official vendor site: qoder.comqoder.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly