GoldieBench deep dive
Qwen 3.8's full benchmark breakdown.
Alibaba's 2.4T flagship — benched through Qoder. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
Provisional flag. Part of this model's run hit infrastructure failures that are being re-run; its average is still settling. Scores shown are real judged builds only.
02 · Every benchmark, bar by bar
All 41 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Gorgeous multi-burst scene with dense sparks, long flowing trails, hue-tinted flashes, and a full city skyline with lit windows plus mirrored water reflections — the teal/magenta glow band and shimmer sell the harbor atmosphere. Strong interactivity (click/tap + auto-fire) and varied explosion types; this is a top-tier fireworks build.”
“Gorgeous, dramatic drape with visible weave texture, corner-pinned peaks and the teal sphere reading clearly through the fabric — cinematic lighting and vignette elevate it above the field; only minor nit is the cloth resolution/collision looks slightly coarse at the sphere contact.”
“Stunning GPU Mandelbrot with smooth iteration coloring, distance-estimation rim shading and a gorgeous editorial HUD (readouts, palettes, tour) — the deep spiral detail renders beautifully. Strong on both polish and interactivity; only nit is the somewhat washed-out interior gradient, but it clearly competes for the top of the field.”
“Strong: a genuinely beautiful, dense spiral with a glowing gold core, color-graded arms (gold→pink→blue) and a polished editorial HUD with live telemetry and interaction hints. Weak: the low 20 fps telemetry hints at heavy load, but the render itself is gorgeous and clearly on-brief, edging past the field's best.”
“Gorgeous rendering — convincing metaball wax with glowing blobs rising in a beautifully detailed brass-and-glass lamp, plus Monoton neon title, live temp readouts, heat slider and theme chips. Polished, on-brief, and clearly at the top of the field; only nitpick is empty left-side negative space.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Qwen 3.8's lowest scored builds, verdicts unedited:
“Polished HUD (Cinzel headers, vigor/breath meters, objective, axe crosshair) and the scene renders a coherent low-poly ruin with columns, arches and torch posts, but the lighting reads flat/dim rather than genuinely 'torch-lit,' the space feels sparse, and crucially no visible enemies appear (0/5 slain) so the combat core is unproven — works but generic against a 9.5 best.”
“Gorgeous atmospheric desert render with polished HUD, custom crosshair, and a full combat/wave framework in code — but the screenshot shows a chaotic black jumble of geometry (likely the player craft rendering wrong/clipped) instead of a readable ship or visible enemies, and hull already at 010 with 0 kills suggests broken presentation. Strong art direction, but the central subject looks glitched rather than shippable.”
“Strong neon aesthetic with polished HUD, CRT vignette, and clean paused panel, but the screenshot shows an empty playfield with no visible snake or food, so the actual gameplay isn't demonstrated. Presentation is above-average yet it reads as a shell rather than a proven, running game.”
“Excellent military HUD polish—compass tape, radar with contacts, hull/thr/boost bars, and a well-modeled aircraft all render cleanly. But the screenshot shows no visible enemies in the sky and zero combat happening (KILLS 00, empty airspace), leaving it feeling more like a polished flight sim than a proven dogfight.”
05 · Category breakdown vs the whole field
Games (18 tasks)
Others (3 tasks)
Pages (2 tasks)
Sims (11 tasks)
Visuals (7 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Qwen 3.8 perform on GoldieBench?
Qwen 3.8 averages 8.22/10 across 41 scored one-shot build tasks, ranking #2 of 18 ranked frontier models, with 10 task golds, 9 silvers and 5 bronzes. Every score is a real judged build you can open and play on this site.
02What is Qwen 3.8 best at?
Its strongest scored build is Fireworks at 9.0/10. By category it averages Game 7.9, Other 7.9, Page 8.4, Sim 8.6, Visual 8.5 on the bench.
03How much does Qwen 3.8 cost?
Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot.
Source ledger
- 01Official vendor site: qoder.comqoder.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.