
Real head-to-head · same prompt, one shot
Qwen 3.8 vs Muse Spark 1.2
Alibaba's 2.4T flagship — benched through Qoder. vs Meta's coding reasoning model — co-trained with its own agent, 1M-token window.
Head-to-head verdict: Qwen 3.8 wins 29–6 with 10 ties.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Qwen 3.8 and Muse Spark 1.2, side by side, on 45 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Qwen 3.8 · Benched on GoldieBench via the Qoder CLI (`qoder-qwen`, model Qwen3.8-Max-Preview) — the only door while it has no public API. Non-game tasks are one-shot like the rest of the field. GAME tasks run in Qoder's real AGENT mode: skill-infused build, then up to 2 QA fix rounds where a vision judge + live console errors are fed back and Qwen 3.8 edits its own file (it took crypt from a black-screen 3.0 to a torch-lit 7.8). Scored on a mid-play frame by the same Opus judge as everyone else.
Muse Spark 1.2 · Cloud coder via OpenRouter; the Muse Code agent (one-command install) is its native harness.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Qwen 3.8
Muse Spark 1.2
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Other
Other
Other
Page
Page
Page
Where Qwen 3.8 beat Muse Spark 1.2
The tasks where I gave Qwen 3.8 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Outrun
Game
Qwen 3.8 8.7
·
Muse Spark 1.2 3.2
(+5.5)
· Neon Combat Outrun
What I saw: Gorgeous synthwave scene nails the brief—striped sun, purple mountains, palm silhouettes, glowing grid floor and a chunky neon car with real pseudo-3D road, plus combat elements (crosshair, kills/hostiles, radar, visible enemy vehicles) that elevate it above empty driving sims. H…
Skyrim
Game
Qwen 3.8 8.6
·
Muse Spark 1.2 4.2
(+4.4)
· Skeletons at dusk
What I saw: Strong first-person fantasy scene with a beautiful dusk palette, low-poly terrain and trees, plus visible skeleton enemies, a held weapon/shield, working HUD (Skyrim-style HP/magicka/stamina), compass, kill quest and DETECTED stealth chip — genuinely combat-ready rather than an e…
Crypt
Game
Qwen 3.8 7.8
·
Muse Spark 1.2 4.5
(+3.3)
What I saw: Strong torch-lit atmosphere with warm flickering light, stone/rune textures, a polished chamfered HUD (health/torch/kills/depth/gold) and a visible skeleton enemy plus first-person weapon — genuinely on-brief. Held back by the oddly framed/oversized weapon dominating the view and…
Dragonflight
Game
Qwen 3.8 7.0
·
Muse Spark 1.2 4.2
(+2.8)
What I saw: Strong polished HUD (vitality/fury meters, score, rings 0/22, kills, crosshair, controls strip) and a nicely modelled dragon with shadow over a moody sunset desert, but the screenshot shows NO neon rings, no visible enemies/combat, and no fire-breath firing — the core on-brief 'n…
Lavalamp
Visual
Qwen 3.8 8.7
·
Muse Spark 1.2 6.5
(+2.2)
· Metaball wax realism
What I saw: Gorgeous rendering — convincing metaball wax with glowing blobs rising in a beautifully detailed brass-and-glass lamp, plus Monoton neon title, live temp readouts, heat slider and theme chips. Polished, on-brief, and clearly at the top of the field; only nitpick is empty left-sid…
Where Muse Spark 1.2 beat Qwen 3.8
The tasks where I gave Muse Spark 1.2 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Waves
Visual
Muse Spark 1.2 6.2
·
Qwen 3.8 2.0
(+4.2)
What I saw: The polished UI (control panel, buoy, islands, ripples) is impressive and clearly on-brief, but the actual ocean surface looks nearly flat and washed-out white with no visible wave undulation, and the 10 FPS readout signals performance trouble — the core 'animated wave simulation…
Game
Game
Muse Spark 1.2 8.4
·
Qwen 3.8 6.5
(+1.9)
· polished 3D shooter
What I saw: Strong: fully rendered Three.js 3D arena shooter with a cohesive premium sci-fi HUD (hull, thruster, score, hostiles), crosshair, and enemy ships/rovers on a lit terrain. Weak: the bright washed-out palette makes low-poly rocks and hero read a bit flat/plain, keeping it just shy …
Dogfight
Game
Muse Spark 1.2 8.4
·
Qwen 3.8 6.8
(+1.6)
What I saw: Strong, polished chase-cam dogfight: crisp HUD with meters/radar/lock indicator, clean 3D plane and volumetric clouds, visible bandits and a lock-on target — very shippable. Slightly held back from top by generic sky-only environment and no visible combat action (tracers/effects)…
Dragonrealm
Game
Muse Spark 1.2 8.3
·
Qwen 3.8 7.8
(+0.5)
What I saw: Strong, polished frozen-world render with cohesive HUD (health/stamina/compass/kills), a nicely stylized low-poly hero, snow terrain, pines and rocks with clean shadows — clearly on-brief and shippable, though it reads more diorama than expansive Skyrim vista and no dragon/enemie…
Webos
Page
Muse Spark 1.2 8.6
·
Qwen 3.8 8.4
(+0.2)
· polished macOS desktop
What I saw: Strong, cohesive macOS-style build with all three apps functional and visible (Notes with toolbar, Paint with working strokes/palette, Terminal responding to ls/echo), plus a Files app, glossy dock, topbar with live clock, and a gorgeous animated gradient wallpaper. Only minor ni…
Strengths & weaknesses I logged
Qwen 3.8
Strengths
- Skyrim-style open worlds — the Dragon Realm build rendered a lit snowfield, first-person sword and working roaming enemies (7.8)
- Held up across game genres early — voxel sandbox and Doom raycaster both came out shippable
- Runs as a real agentic coder inside Qoder (writes + iterates on files), not just a chat model
Trade-offs
- Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift
- Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can
Muse Spark 1.2
Strengths
- Generative art & shader-feel scenes (fractal 8.7, aurora/galaxy/matrix/synthwave 8.6)
- Full app chrome one-shot (macOS-clone desktop 8.6)
- Fast one-shots — most builds landed in 45-80s
- 1M context for whole-repo work
Trade-offs
- 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5)
- Open-world briefs collapse to HUD-only shells
- Reasoning tokens billed as output
Pricing & context — the spec sheet
| Spec | Qwen 3.8 | Muse Spark 1.2 |
|---|---|---|
| Vendor | Alibaba | Meta |
| Context window | Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. | 1,000,000 tokens |
| Price | Qoder plan | $1.25 in / $4.25 out per 1M |
| Pricing detail | Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot. | Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field. |
| Release | 2026-07 | 2026-08-05 |
| Bench coverage | 45/45 scored · avg 8.10/10 | 50/50 scored · avg 7.47/10 |
The verdict — which should you pick?
Across 45 scored shared tasks, Qwen 3.8 averaged 8.10/10, beating Muse Spark 1.2's 7.47/10 by 0.63 points. Pick Qwen 3.8 when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Qwen 3.8 and Muse Spark 1.2 both into the Agent Operating System and dispatch each from the kanban by task type — one-shot 3d game and world prototypes where atmosphere matters → Qwen 3.8, generative-art visuals → Muse Spark 1.2. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Qwen 3.8 vs Muse Spark 1.2
Which is better, Qwen 3.8 or Muse Spark 1.2?
On Goldie Bench, Qwen 3.8 averages 8.10/10 across the shared tasks, with 8 gold, 7 silver, 7 bronze overall. Muse Spark 1.2 averages 7.47/10, with 0 gold, 3 silver, 4 bronze. Qwen 3.8 wins the head-to-head 29–6.
How much does Qwen 3.8 cost vs Muse Spark 1.2?
Qwen 3.8: Qwen3.8-Max-Preview is Alibaba's ~2.4T-parameter flagship, positioned just behind Claude Fable 5. It is NOT on OpenRouter or a public API yet — the only access today is inside Qoder (Alibaba's agentic coding platform, free 2-week Pro trial). Benched here via the Qoder CLI on model `Qwen3.8-Max-Preview`, one-shot. Muse Spark 1.2: Meta's coding-optimized reasoning model, released 2026-08-05 beside the Muse Code agent. $0.15/1M cached input. Contributor tier is token-rate-limited in a rolling 5-hour window. Benched release-day via OpenRouter (meta/muse-spark-1.2, first-party listing); Opus 4.8 judged every real rendered poster, same rubric as the whole field.
What's the context window for Qwen 3.8 vs Muse Spark 1.2?
Qwen 3.8 has a Served through Alibaba's Qoder agent platform; the 3.8-Max preview has no standalone public context window yet. context window. Muse Spark 1.2 has a 1,000,000 tokens context window.
When should I pick Qwen 3.8 over Muse Spark 1.2?
Pick Qwen 3.8 for: One-shot 3D game and world prototypes where atmosphere matters; Anyone already in the Qoder IDE/CLI wanting a near-frontier model free on the Pro trial; A cheaper stand-in for Fable 5 on creative-visual builds. The trade-off is the weaknesses we logged on the bench: Torch-lit dungeon (crypt) came out generic (6.3); one open-world RPG one-shot black-screened (twilightvale 2.5) — classic three.js r128 API drift; Preview is Qoder / Token-Plan only — no OpenRouter or public API, so it can't be routed into an app the way the OpenRouter models can.
When should I pick Muse Spark 1.2 over Qwen 3.8?
Pick Muse Spark 1.2 for: Generative-art visuals; Dashboard & app-shell one-shots; Long-context refactors (1M window). The trade-off is the weaknesses we logged on the bench: 3D game worlds often render black/empty (dragonrealm 2.5, dogfight 3.0, doom 3.5); Open-world briefs collapse to HUD-only shells; Reasoning tokens billed as output.
How does Goldie Bench score Qwen 3.8 vs Muse Spark 1.2?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Qwen 3.8 vs Fusion Muse Spark 1.2 vs Fusion Qwen 3.8 vs Claude Opus 5 Muse Spark 1.2 vs Claude Opus 5 Qwen 3.8 vs Hermes MoA Muse Spark 1.2 vs Hermes MoA Qwen 3.8 vs GPT-5.6 Sol Muse Spark 1.2 vs GPT-5.6 SolFull model pages: Qwen 3.8 · Muse Spark 1.2 · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































