Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Alibaba

Qwen 3.7

Multilingual open-weights — strong on Chinese reasoning.

Context256,000 tokens
PricingOpen weights · free for individuals
Tasks tested47
Avg score7.00/10 average
Medals🥇0 🥈0 🥉0
Release2026-06
Official sitechat.qwen.ai ↗
Official vendor source
Qwen 3.7 is built by Alibaba — see the vendor's own product page, pricing, and docs at chat.qwen.ai.
Visit chat.qwen.ai →

What is Qwen 3.7?

Qwen 3.7 is the Alibaba frontier model with a 256,000 tokens context window, released 2026-06. Tagline: Multilingual open-weights — strong on Chinese reasoning.. Official source: chat.qwen.ai.

Pricing detail. Alibaba's open-weights release — downloadable from Hugging Face, runnable locally or via Alibaba Cloud's free tier for individuals.

How I use it inside the Agent OS. Wired alongside GLM-5.2 in Agent OS for open-weights agent loops where you want vendor diversity.

What I built with Qwen 3.7

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Qwen 3.7 shipped on the bench: 47 one-shot demos across 256,000 tokens of context. Of those, 47 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Open weights, free for individuals — same model class as GLM-5.2
  • Best-of-three on fluid simulation in the Goldie Bench bench
  • Multilingual depth — Chinese reasoning especially strong

Trade-offs

  • Only 5 tasks scored on the bench so far — small sample size
  • Trails GLM-5.2 on cinematic visual builds at similar pricing

Best for

  • Open-weights alternative to GLM-5.2 when you want a different model family
  • Multilingual workloads (Chinese, multi-script content)
  • Fluid and particle simulations

Every demo by Qwen 3.7

47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Qwen 3.7 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
13KB · plays clean · webgl, input
Qwen 3.7 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
14KB · animation runs but no input response · webgl
Qwen 3.7 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
12KB · plays clean · input
Qwen 3.7 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
12KB · plays clean · webgl
Qwen 3.7 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
12KB · plays clean · webgl
Qwen 3.7 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
25KB · plays clean · plain
Qwen 3.7 one-shot build of Voxelcraft — GoldieBench AI benchmark screenshot▶ LIVE
Voxelcraft
Game
14KB · plays clean · webgl, input
Qwen 3.7 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
21KB · plays clean · three
Qwen 3.7 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
21KB · plays clean · three, webgl
Qwen 3.7 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo
Page
13KB · plays clean · plain
Qwen 3.7 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute
Game
12KB · plays clean · three, webgl, input
Qwen 3.7 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim
Game
12KB · plays clean · three, webgl (re-rolled)
Qwen 3.7 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade
Game
The closest test. All three shipped a real, juicy game. Opus's breakout had the most game-feel (particle bursts + live combo). Qwen's neon breakout is clean and vibrant. GLM went its own way with fullscreen asteroids. Genuinely hard to separate.
Qwen 3.7 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
18KB · plays clean · webgl, input
Qwen 3.7 one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight
Game
13KB · plays clean · webgl, input
Qwen 3.7 one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game
Game
25KB · plays clean · plain
Qwen 3.7 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
21KB · plays clean · webgl, audio, input
Qwen 3.7 one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity
Game
8KB · plays clean · webgl, rAF
Qwen 3.7 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
14KB · plays clean · webgl, input
Qwen 3.7 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
16KB · plays clean · webgl
Qwen 3.7 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
11KB · animation runs but no input response · plain
Qwen 3.7 one-shot build of Pool — GoldieBench AI benchmark screenshot▶ LIVE
Pool
Game
16KB · animation runs but no input response · webgl
Qwen 3.7 one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster
Game
0KB · no animation detected on load · TIMEOUT
Qwen 3.7 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
13KB · animation runs but no input response · input
Qwen 3.7 one-shot build of Landing — GoldieBench AI benchmark screenshot▶ LIVE
Landing
Page
GLM and Opus both produced premium gradient 'Intelligence, reimagined / distilled' keynote heroes — basically a tie. Qwen's is clean and well-built (proper nav + three feature cards) but the headline ('Built for the next generation of builders') lands flatter than the gradient heroes.
Qwen 3.7 one-shot build of Webos — GoldieBench AI benchmark screenshot▶ LIVE
Webos
Page
23KB · animation runs but no input response · plain
Qwen 3.7 one-shot build of Blackhole — GoldieBench AI benchmark screenshot▶ LIVE
Blackhole
Sim
5KB · plays clean · webgl, rAF
Qwen 3.7 one-shot build of Boids — GoldieBench AI benchmark screenshot▶ LIVE
Boids
Sim
10KB · plays clean · plain
Qwen 3.7 one-shot build of Cloth — GoldieBench AI benchmark screenshot▶ LIVE
Cloth
Sim
13KB · plays clean · webgl
Qwen 3.7 one-shot build of Fluid — GoldieBench AI benchmark screenshot▶ LIVE
Fluid
Sim
GLM filled the bowl with glowing liquid that genuinely sloshes — the most convincing of the three. Opus's particles glowed but clumped to the centre, and Qwen's is a solid working sim but reads thinner. Play them and tilt — GLM's is the one that feels like fluid.
Qwen 3.7 one-shot build of Fractal — GoldieBench AI benchmark screenshot▶ LIVE
Fractal
Sim
6KB · animation runs but no input response · webgl, input, rAF
Qwen 3.7 one-shot build of Galaxy — GoldieBench AI benchmark screenshot▶ LIVE
Galaxy
Sim
9KB · animation runs but no input response · webgl
Qwen 3.7 one-shot build of Orbit — GoldieBench AI benchmark screenshot▶ LIVE
Orbit
Sim
Opus nailed the brief — distinct labelled planet orbits, a real NEO panel, a sim clock. GLM went dramatic with a glowing nebula swirl (gorgeous, but more galaxy than orbit map). Qwen drew a dense, busy orbital swarm — structurally orbit-like but dimmer and harder to read.
Qwen 3.7 one-shot build of Particleforge — GoldieBench AI benchmark screenshot▶ LIVE
Particleforge
Sim
12KB · plays clean · webgl
Qwen 3.7 one-shot build of Pathtracer — GoldieBench AI benchmark screenshot▶ LIVE
Pathtracer
Sim
12KB · plays clean · webgl
Qwen 3.7 one-shot build of Reactiondiff — GoldieBench AI benchmark screenshot▶ LIVE
Reactiondiff
Sim
14KB · animation runs but no input response · webgl
Qwen 3.7 one-shot build of Solar — GoldieBench AI benchmark screenshot▶ LIVE
Solar
Sim
7KB · plays clean · webgl, controls, rAF
Qwen 3.7 one-shot build of Wormhole — GoldieBench AI benchmark screenshot▶ LIVE
Wormhole
Sim
7KB · plays clean · webgl, input, rAF
Qwen 3.7 one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora
Visual
6KB · plays clean · webgl, rAF
Qwen 3.7 one-shot build of Fireworks — GoldieBench AI benchmark screenshot▶ LIVE
Fireworks
Visual
7KB · plays clean · webgl, rAF
Qwen 3.7 one-shot build of Lavalamp — GoldieBench AI benchmark screenshot▶ LIVE
Lavalamp
Visual
6KB · animation runs but no input response · webgl, rAF
Qwen 3.7 one-shot build of Matrix — GoldieBench AI benchmark screenshot▶ LIVE
Matrix
Visual
2KB · plays clean · plain
Qwen 3.7 one-shot build of Plasma — GoldieBench AI benchmark screenshot▶ LIVE
Plasma
Visual
8KB · plays clean · input, rAF
Qwen 3.7 one-shot build of Synthwave — GoldieBench AI benchmark screenshot▶ LIVE
Synthwave
Visual
9KB · animation runs but no input response · webgl
Qwen 3.7 one-shot build of Terrain — GoldieBench AI benchmark screenshot▶ LIVE
Terrain
Visual
4KB · plays clean · webgl, rAF
Qwen 3.7 one-shot build of Voxel — GoldieBench AI benchmark screenshot▶ LIVE
Voxel
Visual
GLM built the densest, most colourful city (windowed skyscrapers + speed/coins HUD). Opus ran the furthest with the cleanest motion. Qwen's is atmospheric — a foggy tunnel of buildings — but more muted and it crashes quicker.
Qwen 3.7 one-shot build of Waves — GoldieBench AI benchmark screenshot▶ LIVE
Waves
Visual
8KB · plays clean · webgl, rAF
every demo, in a grid · click any one to play

Compare Qwen 3.7 against every other model

Every head-to-head featuring Qwen 3.7. Verdicts shown for scored pairs.

Qwen 3.7 vs Fusion
Fusion leads 46–0
Qwen 3.7 vs Claude Opus 5
Claude Opus 5 leads 41–5
Qwen 3.7 vs Hermes MoA
Hermes MoA leads 42–4
Qwen 3.7 vs GPT-5.6 Sol
GPT-5.6 Sol leads 43–4
Qwen 3.7 vs Claude Fable 5
Claude Fable 5 leads 38–6
Qwen 3.7 vs Qwen 3.8
Qwen 3.8 leads 36–6
Qwen 3.7 vs Grok
Grok leads 34–1
Qwen 3.7 vs MiniMax M3
MiniMax M3 leads 34–4
Qwen 3.7 vs Fugu Ultra
Fugu Ultra leads 34–6
Qwen 3.7 vs Kimi K3
Kimi K3 leads 37–10
Qwen 3.7 vs GLM-5.2
GLM-5.2 leads 29–6
Qwen 3.7 vs Fugu Mini
Fugu Mini leads 29–3
Qwen 3.7 vs Opus 4.8
Opus 4.8 leads 21–11
Qwen 3.7 vs Kimi K2.7
Kimi K2.7 leads 11–7
Qwen 3.7 vs Qwable 5 27B Coder
Qwable 5 27B Coder leads 28–12
Qwen 3.7 vs Gemini 3.6 Flash
Gemini 3.6 Flash leads 30–17
Qwen 3.7 vs Claude Sonnet 5
Claude Sonnet 5 leads 32–15
Qwen 3.7 vs Fugu Ultra 1.1
Fugu Ultra 1.1 leads 12–11
Qwen 3.7 vs Inkling
Qwen 3.7 leads 34–13
Qwen 3.7 vs Agents-A1
Qwen 3.7 leads 32–9
Qwen 3.7 vs Gemma 4 12B · MLX
Qwen 3.7 leads 37–5
Qwen 3.7 vs Laguna XS 2.1
Qwen 3.7 leads 38–3
Qwen 3.7 vs Qwythos 9B
Qwen 3.7 leads 42–0
Qwen 3.7 vs LongCat-2.0
LongCat-2.0 leads 3–0
Qwen 3.7 vs Hy3
Qwen 3.7 leads 5–2
Qwen 3.7 vs Gemma-4 12B Coder
Qwen 3.7 leads 6–0
Qwen 3.7 vs DeepSeek V4 Flash
47 shared tasks · unscored
Qwen 3.7 vs DeepSeek V4 Pro
47 shared tasks · unscored
Qwen 3.7 vs Kimi K2.7 · Fast
47 shared tasks · unscored
Qwen 3.7 vs Kimi K2.7 · No-Think
47 shared tasks · unscored
Qwen 3.7 vs Kimi K2.7 · Quality
47 shared tasks · unscored
Qwen 3.7 vs Ornith 1.0
42 shared tasks · unscored
Qwen 3.7 vs Claude Mythos 5
Reference-only
Qwen 3.7 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Qwen 3.7 vs Fusion Qwen 3.7 vs Claude Opus 5 Qwen 3.7 vs Hermes MoA Qwen 3.7 vs GPT-5.6 Sol Qwen 3.7 vs Claude Fable 5 Qwen 3.7 vs Qwen 3.8 Qwen 3.7 vs Grok Qwen 3.7 vs MiniMax M3 Qwen 3.7 vs Fugu Ultra Qwen 3.7 vs Kimi K3 Qwen 3.7 vs GLM-5.2 Qwen 3.7 vs Fugu Mini Qwen 3.7 vs Opus 4.8 Qwen 3.7 vs Kimi K2.7 Qwen 3.7 vs Qwable 5 27B Coder Qwen 3.7 vs Gemini 3.6 Flash Qwen 3.7 vs Claude Sonnet 5 Qwen 3.7 vs Fugu Ultra 1.1 Qwen 3.7 vs Inkling Qwen 3.7 vs Agents-A1 Qwen 3.7 vs Gemma 4 12B · MLX Qwen 3.7 vs Laguna XS 2.1 Qwen 3.7 vs Qwythos 9B Qwen 3.7 vs LongCat-2.0 Qwen 3.7 vs Hy3 Qwen 3.7 vs Gemma-4 12B Coder

Read more on agentos.guide: /glm-vs-qwen-vs-opus, /qwen-hermes

Qwen 3.7 — frequently asked

What is Qwen 3.7?

Qwen 3.7 is Alibaba's AI model — Multilingual open-weights — strong on Chinese reasoning. It has a 256K tokens context window and was released 2026-06.

How good is Qwen 3.7 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 7.00/10 across 47 scored tasks, with 0 gold, 0 silver and 0 bronze medals.

How much does Qwen 3.7 cost?

Open weights · free for individuals. Alibaba's open-weights release — downloadable from Hugging Face, runnable locally or via Alibaba Cloud's free tier for individuals.

Where can I see Qwen 3.7 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly