Qwen 3.8
Alibaba's 2.4T flagship — benched through Qoder.
What is Qwen 3.8?
Qwen 3.8 is the Alibaba frontier model with a Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. context window, released 2026-07. Tagline: Alibaba's 2.4T flagship — benched through Qoder.. Official source: qoder.com.
Pricing detail. Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot.
How I use it inside the Agent OS. Benched on GoldieBench via the Qoder CLI (`qoder-qwen`, model Qwen3.8-Max-Preview) — the only door while it has no public API. Non-game tasks are one-shot like the rest of the field. GAME tasks run in Qoder's real AGENT mode: skill-infused build, then up to 2 QA fix rounds where a vision judge + live console errors are fed back and Qwen 3.8 edits its own file (it took crypt from a black-screen 3.0 to a torch-lit 7.8). Scored on a mid-play frame by the same Opus judge as everyone else. This is a partial run — the full 50-task card replaces it when the batch completes.
What I built with Qwen 3.8
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Qwen 3.8 shipped on the bench: 41 one-shot demos across Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. of context. Of those, 41 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Skyrim-style open worlds — the Dragon Realm build rendered a lit snowfield, first-person sword and working roaming enemies (7.8)
- Held up across game genres early — voxel sandbox and Doom raycaster both came out shippable
- Runs as a real agentic coder inside Qoder (writes + iterates on files), not just a chat model
Trade-offs
- Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift
- Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can
Best for
- One-shot 3D game and world prototypes where atmosphere matters
- Anyone already in the Qoder IDE/CLI wanting a near-frontier model free on the Pro trial
- A cheaper stand-in for Fable 5 on creative-visual builds
Every benchmark — Qwen 3.8's full scorecard
All 41 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Qwen 3.8 deep dive →.
Every demo by Qwen 3.8
41 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Qwen 3.8 against every other model
Every head-to-head featuring Qwen 3.8. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Qwen 3.8 vs Fusion Qwen 3.8 vs Hermes MoA Qwen 3.8 vs GPT-5.6 Sol Qwen 3.8 vs Claude Fable 5 Qwen 3.8 vs Grok Qwen 3.8 vs MiniMax M3 Qwen 3.8 vs Fugu Ultra Qwen 3.8 vs Kimi K3 Qwen 3.8 vs GLM-5.2 Qwen 3.8 vs Fugu Mini Qwen 3.8 vs Opus 4.8 Qwen 3.8 vs Kimi K2.7 Qwen 3.8 vs Qwable 5 27B Coder Qwen 3.8 vs Claude Sonnet 5 Qwen 3.8 vs Qwen 3.7 Qwen 3.8 vs Inkling Qwen 3.8 vs Agents-A1 Qwen 3.8 vs Gemma 4 12B · MLX Qwen 3.8 vs Laguna XS 2.1 Qwen 3.8 vs Qwythos 9B Qwen 3.8 vs LongCat-2.0 Qwen 3.8 vs Hy3 Qwen 3.8 vs Gemma-4 12B CoderRead more on agentos.guide:
Qwen 3.8 — frequently asked
What is Qwen 3.8?
Qwen 3.8 is Alibaba's AI model — Alibaba's 2.4T flagship — benched through Qoder. It has a Qoder-hosted context window and was released 2026-07.
How good is Qwen 3.8 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 8.22/10 across 41 scored tasks, with 10 gold, 9 silver and 5 bronze medals.
How much does Qwen 3.8 cost?
Qoder plan. Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, fr
Where can I see Qwen 3.8 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.