Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Alibaba

Qwen 3.8

Alibaba's 2.4T flagship — benched through Qoder.

ContextServed through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet.
PricingQoder plan
Tasks tested41
Avg score8.22/10 average
Medals🥇10 🥈9 🥉5
Release2026-07
Official siteqoder.com ↗
Official vendor source
Qwen 3.8 is built by Alibaba — see the vendor's own product page, pricing, and docs at qoder.com.
Visit qoder.com →

What is Qwen 3.8?

Qwen 3.8 is the Alibaba frontier model with a Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. context window, released 2026-07. Tagline: Alibaba's 2.4T flagship — benched through Qoder.. Official source: qoder.com.

Pricing detail. Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot.

How I use it inside the Agent OS. Benched on GoldieBench via the Qoder CLI (`qoder-qwen`, model Qwen3.8-Max-Preview) — the only door while it has no public API. Non-game tasks are one-shot like the rest of the field. GAME tasks run in Qoder's real AGENT mode: skill-infused build, then up to 2 QA fix rounds where a vision judge + live console errors are fed back and Qwen 3.8 edits its own file (it took crypt from a black-screen 3.0 to a torch-lit 7.8). Scored on a mid-play frame by the same Opus judge as everyone else. This is a partial run — the full 50-task card replaces it when the batch completes.

What I built with Qwen 3.8

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Qwen 3.8 shipped on the bench: 41 one-shot demos across Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. of context. Of those, 41 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Skyrim-style open worlds — the Dragon Realm build rendered a lit snowfield, first-person sword and working roaming enemies (7.8)
  • Held up across game genres early — voxel sandbox and Doom raycaster both came out shippable
  • Runs as a real agentic coder inside Qoder (writes + iterates on files), not just a chat model

Trade-offs

  • Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift
  • Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can

Best for

  • One-shot 3D game and world prototypes where atmosphere matters
  • Anyone already in the Qoder IDE/CLI wanting a near-frontier model free on the Pro trial
  • A cheaper stand-in for Fable 5 on creative-visual builds

Every demo by Qwen 3.8

41 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Qwen 3.8 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
Strong torch-lit atmosphere with warm flickering light, stone/rune textures, a polished chamfered HUD (health/torch/kills/depth/gold) and a visible skeleton enemy plus first-person weapon — genuinely on-brief. Held back by the oddly framed/oversized weapon dominating the view and a cramped composition that reads more cluttered than commanding, so it's shippable but not quite a task-topper.
Qwen 3.8 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
Strong atmospheric frozen world with snow, pines, mountains, ruins, frozen pond and a working combat loop (visible enemies, KILLS 3/5, 'DRAGR SLAIN!' feedback and HP/stamina HUD). Weak points: the clustered enemy blobs look crude and overlapping, and the sword/character presentation reads more prototype than the polished Skyrim vibe the brief wants, keeping it shippable but short of the field's best.
Qwen 3.8 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
Strong HUD polish (demons-left counter, kills 2/7, minimap, shells, weapon sprite) and textured raycasting all render cleanly, and the source has full demon sprites with chase/attack AI. Weakest point is the screenshot shows no visible enemy on-screen and the fog/lighting is oddly washed-out bright, so combat presence is inferred from code rather than seen.
Qwen 3.8 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing 🥉
Game
Strong render: sleek third-person hovercraft with a banked track, spike-mine hostiles clustered ahead, tire stacks/pillar obstacles, and a gorgeously cohesive sunset HUD with minimap, crosshair, and combat stats. Beats generic walking sims by delivering visible enemies + combat framing; only slight worry is enemies read as static mines rather than aggressive AI, but the polish and on-brief execution top the field.
Qwen 3.8 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim 🥉
Game
Strong first-person fantasy scene with a beautiful dusk palette, low-poly terrain and trees, plus visible skeleton enemies, a held weapon/shield, working HUD (Skyrim-style HP/magicka/stamina), compass, kill quest and DETECTED stealth chip — genuinely combat-ready rather than an empty walking sim. Only minor knocks: enemies look passive/static in the shot and the aesthetic reads more pastel-storybook than gritty Skyrim, but polish and on-brief execution are top-tier.
Qwen 3.8 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
Strong twilight atmosphere with cohesive low-poly village, snow, lit lanterns, character model, and polished HUD (minimap, quest, elite bar) all rendering cleanly on-brief. Held just under top tier because the screenshot shows no visible enemies engaged and combat/weather can't be confirmed in the still, though the framework and elite/kill systems are clearly present.
Qwen 3.8 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot 🥉
Game
Strong dusk-lit third-person city with polished HUD (vitals, ammo, wanted stars, minimap, takedowns), a proper armed player character, pedestrians, park, crosswalks, and cars — clearly on-brief and shippable. Held just below the top tier because the screenshot shows peaceful pedestrians rather than visible combat/enemies engaged, so the shooting-vs-hostiles loop isn't demonstrably firing on screen.
Qwen 3.8 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive 🥈
Game
Gorgeous dusk-lit voxel city with a polished HUD (stars, health/car bars, minimap, cash, cops-down) and visible traffic plus a well-modeled player car with drive-by crosshair; the atmospheric skybox, shadows, and complete GTA feature set (jack cars, wanted stars, cop combat) push it to the top of the field.
Qwen 3.8 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo 🥇
Page
The cold-open scene renders beautifully — massive Bebas Neue headline with gold 'NO OFF HOURS' accent, typed monospace line, HUD framing, film grain and timeline bar all read as a polished Remotion-style reel; source confirms a full 4-scene timeline with animated count-up stats, montage chips and end CTA. Loses a touch since only scene 1 is visible and stats/CTA polish can't be verified in-frame, but the craft and on-brief cinematic direction clearly rival the field's best.
Qwen 3.8 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim 🥈
Game
Gorgeous dawn-lit terrain with runway, hangar, full HUD (IAS/ALT/VS/heading tape/attitude indicator/throttle/hull/minimap) and a takeoff prompt actively firing at 43kt — clearly on-brief and polished with visible birds as targets. Falls just short of the field's best due to the slightly toy-like aircraft model and unproven landing loop, but a strong, shippable flightsim.
Qwen 3.8 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade
Game
Ambitious 3D reinterpretation of Snake with genuine combat (visible enemies, threats counter, hull/boost systems) and a highly polished neon HUD; strong presentation but it drifts from the classic arcade brief into a bespoke shooter-serpent hybrid, and the scene reads a bit cluttered rather than crisp.
Qwen 3.8 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
Excellent military HUD polish—compass tape, radar with contacts, hull/thr/boost bars, and a well-modeled aircraft all render cleanly. But the screenshot shows no visible enemies in the sky and zero combat happening (KILLS 00, empty airspace), leaving it feeling more like a polished flight sim than a proven dogfight.
Qwen 3.8 one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight
Game
Strong polished HUD (vitality/fury meters, score, rings 0/22, kills, crosshair, controls strip) and a nicely modelled dragon with shadow over a moody sunset desert, but the screenshot shows NO neon rings, no visible enemies/combat, and no fire-breath firing — the core on-brief 'neon rings at speed' element is absent from view, making it feel like a walking sim.
Qwen 3.8 one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game
Game
Gorgeous atmospheric desert render with polished HUD, custom crosshair, and a full combat/wave framework in code — but the screenshot shows a chaotic black jumble of geometry (likely the player craft rendering wrong/clipped) instead of a readable ship or visible enemies, and hull already at 010 with 0 kills suggests broken presentation. Strong art direction, but the central subject looks glitched rather than shippable.
Qwen 3.8 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster 🥈
Game
Stunning outrun-style 3D perspective with pyramids, sun, grid floor and a beautifully rendered ship plus polished HUD, wave banner, and enemies inbound — visually well above the field's typical 2D shooters. Screenshot shows active wave (score 5250, 7 kills, hostiles visible), and the source confirms full waves/bosses/power-ups/screen-shake/synth systems, making it a strong task winner.
Qwen 3.8 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
Strong on-brief render: glowing vapor trails, neon grid floor, cyberpunk skyline and a polished HUD with hull/boost/wave systems plus visible red enemy craft engaged in combat. Slightly noisy vehicle silhouette and enemies clustered oddly against the wall keep it just below the top mark, but it clearly delivers the racer + particle + combat brief.
Qwen 3.8 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
Polished HUD (Cinzel headers, vigor/breath meters, objective, axe crosshair) and the scene renders a coherent low-poly ruin with columns, arches and torch posts, but the lighting reads flat/dim rather than genuinely 'torch-lit,' the space feels sparse, and crucially no visible enemies appear (0/5 slain) so the combat core is unproven — works but generic against a 9.5 best.
Qwen 3.8 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun 🥇
Game
Gorgeous synthwave scene nails the brief—striped sun, purple mountains, palm silhouettes, glowing grid floor and a chunky neon car with real pseudo-3D road, plus combat elements (crosshair, kills/hostiles, radar, visible enemy vehicles) that elevate it above empty driving sims. HUD is exceptionally polished; only nitpick is the road geometry looking slightly flat/skewed, but overall this competes for the top of the field.
Qwen 3.8 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
Strong, polished 3D top-down action-RPG with a gorgeous low-poly world, working enemies (threat 10, active combat with player being surrounded), crisp HUD/minimap and ember-collection objective. Combat and sprites deliver, but it leans more twin-stick arena than full RPG (inventory not visibly demonstrated), keeping it just below the task's best.
Qwen 3.8 one-shot build of Matrixrain — GoldieBench AI benchmark screenshot▶ LIVE
Matrixrain 🥈
Other
Strong Katakana rain with convincing depth layers, bright leaders, red anomalies and a visible click-shockwave rippling through the glyphs, all wrapped in a polished CRT HUD with themes, decoding tagline and live stats. Minor generic-topic ceiling, but the parallax, shockwave interaction and typography push it above the field's best.
Qwen 3.8 one-shot build of Mlx Speedtest — GoldieBench AI benchmark screenshot▶ LIVE
Mlx Speedtest 🥇
Other
Strong on-brief render: crisp MLX branding, 6 discovered models with tok/s estimates, phased benchmark steps, a 3D gauge with needle, live log, platform specs and a throughput chart — highly polished and cohesive. Slight weakness is the paused/loading state means the gauge reads 0 and results table is empty in this shot, but the simulation architecture is clearly complete and ships above the field.
Qwen 3.8 one-shot build of Neonsnake — GoldieBench AI benchmark screenshot▶ LIVE
Neonsnake 🥉
Other
Strong neon aesthetic with polished HUD, CRT vignette, and clean paused panel, but the screenshot shows an empty playfield with no visible snake or food, so the actual gameplay isn't demonstrated. Presentation is above-average yet it reads as a shell rather than a proven, running game.
Qwen 3.8 one-shot build of Landing — GoldieBench AI benchmark screenshot▶ LIVE
Landing
Page
Strong cohesive freight/logistics theme with polished dark-teal palette, mono typography, animated departures board (in source) and a confident gradient headline; the visible screenshot is a features section rather than the hero, and the faint globe/route background reads a touch empty, keeping it just shy of the top.
Qwen 3.8 one-shot build of Blackhole — GoldieBench AI benchmark screenshot▶ LIVE
Blackhole
Sim
Strong Interstellar-style render: bright photon ring hugging the black silhouette with the accretion disk arcing over the top and secondary lower image reads convincingly as lensing. Slightly let down by the scattered clumpy orbiting particles that look noisy against the clean disk, keeping it just shy of the top mark.
Qwen 3.8 one-shot build of Boids — GoldieBench AI benchmark screenshot▶ LIVE
Boids 🥇
Sim
Strong 3D boids with full separation/alignment/cohesion, cone birds, trails, bloom post-processing, and a polished neon-on-dark aesthetic that reads immediately as an emergent flock; only minor knock is the still frame can't confirm cohesion feels less clumped mid-scatter, but the render is gorgeous and clearly on-brief — a task winner.
Qwen 3.8 one-shot build of Cloth — GoldieBench AI benchmark screenshot▶ LIVE
Cloth 🥇
Sim
Gorgeous, dramatic drape with visible weave texture, corner-pinned peaks and the teal sphere reading clearly through the fabric — cinematic lighting and vignette elevate it above the field; only minor nit is the cloth resolution/collision looks slightly coarse at the sphere contact.
Qwen 3.8 one-shot build of Fluid — GoldieBench AI benchmark screenshot▶ LIVE
Fluid
Sim
Gorgeous glowing galactic swirl with a molten pulsing core, rich teal/magenta/plum palette and convincing spiral-arm structure driven by curl noise plus interactive mouse stirring; polished and clearly on-brief, though the blown-out white center slightly overwhelms the fluid detail and it reads more nebula than true fluid, keeping it just under the top.
Qwen 3.8 one-shot build of Fractal — GoldieBench AI benchmark screenshot▶ LIVE
Fractal 🥈
Sim
Stunning GPU Mandelbrot with smooth iteration coloring, distance-estimation rim shading and a gorgeous editorial HUD (readouts, palettes, tour) — the deep spiral detail renders beautifully. Strong on both polish and interactivity; only nit is the somewhat washed-out interior gradient, but it clearly competes for the top of the field.
Qwen 3.8 one-shot build of Galaxy — GoldieBench AI benchmark screenshot▶ LIVE
Galaxy 🥇
Sim
Strong: a genuinely beautiful, dense spiral with a glowing gold core, color-graded arms (gold→pink→blue) and a polished editorial HUD with live telemetry and interaction hints. Weak: the low 20 fps telemetry hints at heavy load, but the render itself is gorgeous and clearly on-brief, edging past the field's best.
Qwen 3.8 one-shot build of Orbit — GoldieBench AI benchmark screenshot▶ LIVE
Orbit 🥈
Sim
Stunning render — glowing central star with 371 orbiting bodies, colorful nebulae, and dust lanes create a genuinely cinematic gravitational field, backed by real N-body physics with softening and live telemetry. Polished 'Flight Deck' controls (gravity/time sliders, trails, burst, camera orbit/zoom) make it interactive and shippable; only minor concern is verifying full-field stability, but visually and functionally it tops the task.
Qwen 3.8 one-shot build of Particleforge — GoldieBench AI benchmark screenshot▶ LIVE
Particleforge 🥈
Sim
The rendered galaxy disc is genuinely beautiful — dense purple-to-orange particle field with a glowing molten core, gravity reticle, and rich telemetry/formation HUD that nails the sculpting brief. Held just below top tier by HUD overlap on the left (telemetry/controls panels colliding and clipping the title) which reads as a layout bug in the screenshot.
Qwen 3.8 one-shot build of Pathtracer — GoldieBench AI benchmark screenshot▶ LIVE
Pathtracer 🥈
Sim
Strong: genuine WebGL path tracer with convincing metal/dielectric/lambert spheres, reflections, refraction and soft shadows on a checker floor, plus a polished amber HUD with live SPP/bounce controls. Weak: Monte Carlo noise is still quite grainy and the sky washes to near-flat white, but the physical correctness and interactivity edge the field's best.
Qwen 3.8 one-shot build of Reactiondiff — GoldieBench AI benchmark screenshot▶ LIVE
Reactiondiff 🥇
Sim
Strong Gray-Scott simulation producing a crisp, organic labyrinth pattern with beautiful bioluminescent glow, backed by a polished HUD, 8 presets, live chemistry sliders, and multiple palettes. The rendered field is genuinely on-brief and visually superior to a flat greyscale sim — this tops the field.
Qwen 3.8 one-shot build of Solar — GoldieBench AI benchmark screenshot▶ LIVE
Solar 🥉
Sim
Strong render: convincing textured sun with corona, distinct banded Jupiter, textured Earth/Venus/Mercury, orbit rings, asteroid belt, and a highly polished editorial UI with working clock, info panel, chips, and speed control. Minor nit — the info panel's dashes for sun's orbital year/dist reveal empty fields, but overall this matches the top of the field.
Qwen 3.8 one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora 🥈
Visual
Gorgeous multi-ribbon aurora with convincing teal-to-violet gradients, glowing hems, vertical light shafts, twinkling stars and layered mountain ridges — polished and clearly on-brief. Only minor nit is the slightly stark ridge silhouettes, but the overall composition and depth top the field.
Qwen 3.8 one-shot build of Fireworks — GoldieBench AI benchmark screenshot▶ LIVE
Fireworks 🥇
Visual
Gorgeous multi-burst scene with dense sparks, long flowing trails, hue-tinted flashes, and a full city skyline with lit windows plus mirrored water reflections — the teal/magenta glow band and shimmer sell the harbor atmosphere. Strong interactivity (click/tap + auto-fire) and varied explosion types; this is a top-tier fireworks build.
Qwen 3.8 one-shot build of Lavalamp — GoldieBench AI benchmark screenshot▶ LIVE
Lavalamp 🥇
Visual
Gorgeous rendering — convincing metaball wax with glowing blobs rising in a beautifully detailed brass-and-glass lamp, plus Monoton neon title, live temp readouts, heat slider and theme chips. Polished, on-brief, and clearly at the top of the field; only nitpick is empty left-side negative space.
Qwen 3.8 one-shot build of Matrix — GoldieBench AI benchmark screenshot▶ LIVE
Matrix 🥇
Visual
Gorgeous multi-layer depth with blurred far field, white-hot heads, katakana glyphs and an EMP pulse ring — plus polished CRT scanlines, glitch title, themes and full HUD controls. Amber-tinted mutation streaks and rich feature set push it to the top of the field.
Qwen 3.8 one-shot build of Plasma — GoldieBench AI benchmark screenshot▶ LIVE
Plasma
Visual
Gorgeous multi-oscillator plasma with rich bloom particles, polished HUD, six labeled palette chips with swatches, and full controls (click ripples, keyboard palette, pause) — clearly on-brief and shippable, though the low 9 FPS reading and the paused 'STANDBY' state raise a slight performance concern versus the top of the field.
Qwen 3.8 one-shot build of Synthwave — GoldieBench AI benchmark screenshot▶ LIVE
Synthwave
Visual
Gorgeous striped sun, layered magenta mountains, floating wireframe polyhedra and a crisp CRT-scanline grid make this a textbook synthwave scene with strong polish; the full control panel (themes, speed, sound, pause, BPM/REC HUD) elevates it above the field, only slightly held back by the grid perspective reading a touch flat near the horizon.
Qwen 3.8 one-shot build of Terrain — GoldieBench AI benchmark screenshot▶ LIVE
Terrain
Visual
Strong HUD polish with amber/mono editorial styling, live readouts, and a coherent island terrain with color-banded biomes and floating rocks; but the wireframe-on view reads as a busy dense mesh rather than a striking solid landscape, and the terrain silhouette is fairly flat/low-relief, keeping it just below the top tier.
every demo, in a grid · click any one to play

Compare Qwen 3.8 against every other model

Every head-to-head featuring Qwen 3.8. Verdicts shown for scored pairs.

Qwen 3.8 vs Fusion
Fusion leads 24–13
Qwen 3.8 vs Hermes MoA
Qwen 3.8 leads 19–9
Qwen 3.8 vs GPT-5.6 Sol
Tied 17–17
Qwen 3.8 vs Claude Fable 5
Qwen 3.8 leads 20–12
Qwen 3.8 vs Grok
Qwen 3.8 leads 26–8
Qwen 3.8 vs MiniMax M3
Qwen 3.8 leads 28–10
Qwen 3.8 vs Fugu Ultra
Qwen 3.8 leads 24–9
Qwen 3.8 vs Kimi K3
Qwen 3.8 leads 22–15
Qwen 3.8 vs GLM-5.2
Qwen 3.8 leads 28–10
Qwen 3.8 vs Fugu Mini
Qwen 3.8 leads 23–7
Qwen 3.8 vs Opus 4.8
Qwen 3.8 leads 30–8
Qwen 3.8 vs Kimi K2.7
Qwen 3.8 leads 16–4
Qwen 3.8 vs Qwable 5 27B Coder
Qwen 3.8 leads 25–4
Qwen 3.8 vs Claude Sonnet 5
Qwen 3.8 leads 27–8
Qwen 3.8 vs Qwen 3.7
Qwen 3.8 leads 33–5
Qwen 3.8 vs Inkling
Qwen 3.8 leads 38–2
Qwen 3.8 vs Agents-A1
Qwen 3.8 leads 35–1
Qwen 3.8 vs Gemma 4 12B · MLX
Qwen 3.8 leads 33–1
Qwen 3.8 vs Laguna XS 2.1
Qwen 3.8 leads 34–0
Qwen 3.8 vs Qwythos 9B
Qwen 3.8 leads 33–1
Qwen 3.8 vs LongCat-2.0
LongCat-2.0 leads 2–1
Qwen 3.8 vs Hy3
Qwen 3.8 leads 6–0
Qwen 3.8 vs Gemma-4 12B Coder
Qwen 3.8 leads 6–0
Qwen 3.8 vs Kimi K2.7 · Fast
38 shared tasks · unscored
Qwen 3.8 vs Kimi K2.7 · No-Think
38 shared tasks · unscored
Qwen 3.8 vs Kimi K2.7 · Quality
38 shared tasks · unscored
Qwen 3.8 vs Ornith 1.0
34 shared tasks · unscored
Qwen 3.8 vs Claude Mythos 5
Reference-only
Qwen 3.8 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Qwen 3.8 vs Fusion Qwen 3.8 vs Hermes MoA Qwen 3.8 vs GPT-5.6 Sol Qwen 3.8 vs Claude Fable 5 Qwen 3.8 vs Grok Qwen 3.8 vs MiniMax M3 Qwen 3.8 vs Fugu Ultra Qwen 3.8 vs Kimi K3 Qwen 3.8 vs GLM-5.2 Qwen 3.8 vs Fugu Mini Qwen 3.8 vs Opus 4.8 Qwen 3.8 vs Kimi K2.7 Qwen 3.8 vs Qwable 5 27B Coder Qwen 3.8 vs Claude Sonnet 5 Qwen 3.8 vs Qwen 3.7 Qwen 3.8 vs Inkling Qwen 3.8 vs Agents-A1 Qwen 3.8 vs Gemma 4 12B · MLX Qwen 3.8 vs Laguna XS 2.1 Qwen 3.8 vs Qwythos 9B Qwen 3.8 vs LongCat-2.0 Qwen 3.8 vs Hy3 Qwen 3.8 vs Gemma-4 12B Coder

Read more on agentos.guide:

Qwen 3.8 — frequently asked

What is Qwen 3.8?

Qwen 3.8 is Alibaba's AI model — Alibaba's 2.4T flagship — benched through Qoder. It has a Qoder-hosted context window and was released 2026-07.

How good is Qwen 3.8 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 8.22/10 across 41 scored tasks, with 10 gold, 9 silver and 5 bronze medals.

How much does Qwen 3.8 cost?

Qoder plan. Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, fr

Where can I see Qwen 3.8 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly