Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Zhipu / Z.ai

GLM-5.2

The never-forgets agent — 1M context, open weights.

Context1,000,000 tokens
PricingOpen weights · free for individuals
Tasks tested47
Avg score7.77/10 average
Medals🥇5 🥈0 🥉0
Release2026-06-14
Official sitez.ai ↗
Official vendor source
GLM-5.2 is built by Zhipu / Z.ai — see the vendor's own product page, pricing, and docs at z.ai.
Visit z.ai →

Reference benchmarks for GLM-5.2

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for GLM-5.2 is honest about what's measured.

SWE-bench Pro
62.1%
vs GLM-5.1 on SWE-bench Pro
+3.7 pts (was 58.4%)
Context window
1M tokens
Cost vs GPT-5.5
~1/6 (~$1.40 in / $4.40 out)
Headline
Beats GPT-5.5 on multiple long-horizon coding benchmarks

What is GLM-5.2?

GLM-5.2 is the Zhipu / Z.ai frontier model with a 1,000,000 tokens context window, released 2026-06-14. Tagline: The never-forgets agent — 1M context, open weights.. Official source: z.ai.

Pricing detail. Open-weights release: weights downloadable from Hugging Face for self-hosting, or runnable for free on z.ai for individuals (commercial use has separate licensing).

How I use it inside the Agent OS. Default model inside Agent OS for any task that touches a long context — codebase Q&A, multi-file refactors, agent memory replay.

What I built with GLM-5.2

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what GLM-5.2 shipped on the bench: 47 one-shot demos across 1,000,000 tokens of context. Of those, 47 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • 1M-token context window — best-in-class long-document and large-codebase work
  • Open weights — runs locally, no vendor lock-in, no token meter
  • Top of the bench for cinematic visuals (neon city, synthwave, voxel runner)

Trade-offs

  • Faceplanted on the Goldie Bench raycaster — the engine was great but it spawned the player inside a wall
  • First-shot reliability lags Opus by a hair on consistency

Best for

  • Long-context agent loops — pasting a whole codebase into one prompt
  • Cinematic visual builds — landing pages, voxel scenes, synthwave runners
  • Anyone who needs to run a frontier coder locally for $0

Every demo by GLM-5.2

47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

GLM-5.2 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
29KB · plays clean · three, webgl, pointer-lock
GLM-5.2 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
142KB · plays clean · plain
GLM-5.2 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
All three are real, playable shooters. Opus drops you in a corridor with an imp dead ahead — gun, crosshair and HUD framed like a screenshot. Kimi matches it: a monster down a textured hall, health, ammo, minimap. GLM ships a gorgeous 'HAZARD PROTOCOL' title screen with a working game behind it, though it too spawns facing a wall. Opus by a hair on the cleanest fight.
GLM-5.2 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
34KB · plays clean · plain
GLM-5.2 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
22KB · plays clean · three, webgl, pointer-lock
GLM-5.2 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
54KB · plays clean · plain
GLM-5.2 one-shot build of Voxelcraft — GoldieBench AI benchmark screenshot▶ LIVE
Voxelcraft
Game
44KB · plays clean · three, webgl
GLM-5.2 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
43KB · plays clean · plain (re-rolled)
GLM-5.2 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
47KB · plays clean · three, webgl
GLM-5.2 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo
Page
25KB · plays clean · plain
GLM-5.2 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute
Game
32KB · plays clean · plain
GLM-5.2 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim
Game
66KB · plays clean · three, webgl
GLM-5.2 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade
Game
All three shipped a genuinely juicy game. Opus's breakout had the most game-feel — particle bursts and a live combo. Kimi's breakout was clean and solid. GLM went its own way with fullscreen neon asteroids. The closest of the practical five.
GLM-5.2 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
43KB · plays clean · webgl
GLM-5.2 one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight
Game
57KB · plays clean · plain
GLM-5.2 one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game
Game
25KB · plays clean · audio
GLM-5.2 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
75KB · plays clean · plain
GLM-5.2 one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity 🥇
Game
GLM's is the most cinematic — neon towers, a setting sun, Japanese signage and a flight HUD, like a frame from a film. Opus's is a clean canyon of lit skyscrapers racing to a vanishing point. Kimi leaned into the synthwave sun and grid more than the city itself. GLM wins the skyline.
GLM-5.2 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
39KB · plays clean · plain
GLM-5.2 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
30KB · plays clean · three, webgl
GLM-5.2 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
GLM shipped the full arcade package — an 'OUTRUN 2086' title, gear, RPM and velocity dials, mountains, the car cruising at 90+. Opus's road curves hard past rumble strips and palms into a scanline sun. Kimi's 'NEON OUTRUN' is clean and on-brief. GLM edges it on sheer completeness.
GLM-5.2 one-shot build of Pool — GoldieBench AI benchmark screenshot▶ LIVE
Pool
Game
46KB · plays clean · plain
GLM-5.2 one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster
Game
Kimi nailed it — brick walls, a checkered floor, a clean minimap, textbook Wolfenstein, runs clean out of the box. Opus's is close and more atmospheric: warm fog and a vignette down a stone corridor (A/D to turn, W/S to move). GLM's engine is genuinely good — brick and mossy-stone walls, fog, a minimap — but its one-shot spawned the player buried inside a wall, dead on arrival; I nudged the start one cell so you can actually walk it. That spawn bug is why it scores lowest here, even though the e
GLM-5.2 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
54KB · plays clean · plain
GLM-5.2 one-shot build of Landing — GoldieBench AI benchmark screenshot▶ LIVE
Landing 🥇
Page
Funniest result of the lot: GLM and Opus independently produced near-identical premium 'Introducing Nova 1 — Intelligence, reimagined / distilled' keynote pages — gradient hero, full nav, pricing tiers. A dead heat. Kimi's was a plainer set of feature cards.
GLM-5.2 one-shot build of Webos — GoldieBench AI benchmark screenshot▶ LIVE
Webos
Page
61KB · plays clean · plain
GLM-5.2 one-shot build of Blackhole — GoldieBench AI benchmark screenshot▶ LIVE
Blackhole
Sim
Opus nailed it — a pure-black event horizon, a bright photon ring, and the disk bent up and over the top exactly like the film's lensing. GLM came in strong with a clean ring and a starfield warping past the hole. Kimi's disk is fine, but the background is a soft grey blur instead of stars. This one's Opus's.
GLM-5.2 one-shot build of Boids — GoldieBench AI benchmark screenshot▶ LIVE
Boids
Sim
24KB · plays clean · plain
GLM-5.2 one-shot build of Cloth — GoldieBench AI benchmark screenshot▶ LIVE
Cloth
Sim
20KB · plays clean · plain
GLM-5.2 one-shot build of Fluid — GoldieBench AI benchmark screenshot▶ LIVE
Fluid 🥇
Sim
GLM filled the bowl with glowing liquid that actually sloshes — the most convincing 'liquid in a bowl'. Opus's particles glowed but clumped to the centre. Kimi's collapsed into a tiny blob.
GLM-5.2 one-shot build of Fractal — GoldieBench AI benchmark screenshot▶ LIVE
Fractal
Sim
All three are genuinely good. Kimi's is the jaw-dropper — a deep rainbow plunge into a seahorse spiral, dense with self-similar detail. Opus zooms smoothly into the seahorse valley with a tasteful cycling palette. GLM frames the whole iconic set in a fire palette with a live coordinate HUD, then descends. Kimi takes this one on raw spectacle.
GLM-5.2 one-shot build of Galaxy — GoldieBench AI benchmark screenshot▶ LIVE
Galaxy
Sim
Opus built a proper interactive 3D galaxy — drag to orbit a 7,000-star cloud around a glowing core. Kimi's is the prettiest single frame: a clean tilted spiral disk with rainbow arms. GLM's runs on a canvas with a slick NGC-style HUD and zoom, just less dramatic at a glance. Three good galaxies, three different bets.
GLM-5.2 one-shot build of Orbit — GoldieBench AI benchmark screenshot▶ LIVE
Orbit
Sim
Opus nailed the brief — labelled planet orbits, a real NEO / close-pass panel, a sim clock. GLM went for drama: a glowing nebula swirl that's gorgeous but reads more galaxy than orbit map. Kimi's is accurate but dim and sparse.
GLM-5.2 one-shot build of Particleforge — GoldieBench AI benchmark screenshot▶ LIVE
Particleforge
Sim
32KB · plays clean · plain
GLM-5.2 one-shot build of Pathtracer — GoldieBench AI benchmark screenshot▶ LIVE
Pathtracer
Sim
28KB · plays clean · webgl
GLM-5.2 one-shot build of Reactiondiff — GoldieBench AI benchmark screenshot▶ LIVE
Reactiondiff
Sim
29KB · plays clean · plain
GLM-5.2 one-shot build of Solar — GoldieBench AI benchmark screenshot▶ LIVE
Solar
Sim
Three genuinely good space sims. Opus tilts the orbits into real 3D with a bloom-heavy sun and Saturn's rings. GLM's is the most product-like — labelled planets, orbit and label toggles, a clean HUD. Kimi's is a tidy tilted-orbit system with rings and a deep starfield. Opus and GLM are neck-and-neck; Opus takes it on the 3D feel.
GLM-5.2 one-shot build of Wormhole — GoldieBench AI benchmark screenshot▶ LIVE
Wormhole
Sim
26KB · plays clean · plain
GLM-5.2 one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora
Visual
8KB · plays clean · plain
GLM-5.2 one-shot build of Fireworks — GoldieBench AI benchmark screenshot▶ LIVE
Fireworks
Visual
16KB · plays clean · plain
GLM-5.2 one-shot build of Lavalamp — GoldieBench AI benchmark screenshot▶ LIVE
Lavalamp
Visual
6KB · plays clean · three, webgl, rAF
GLM-5.2 one-shot build of Matrix — GoldieBench AI benchmark screenshot▶ LIVE
Matrix
Visual
6KB · plays clean · rAF
GLM-5.2 one-shot build of Plasma — GoldieBench AI benchmark screenshot▶ LIVE
Plasma
Visual
27KB · plays clean · webgl
GLM-5.2 one-shot build of Synthwave — GoldieBench AI benchmark screenshot▶ LIVE
Synthwave 🥇
Visual
This is GLM's. A cyan wireframe mountain range scrolling under a scanline synthwave sun — the single most beautiful frame in the whole shoot-out. Opus's clean Tron grid and magenta horizon is a close, cooler-toned second. Kimi got the idea but blew the exposure — the grid washes out to near-white. GLM wins this one going away.
GLM-5.2 one-shot build of Terrain — GoldieBench AI benchmark screenshot▶ LIVE
Terrain
Visual
13KB · plays clean · plain
GLM-5.2 one-shot build of Voxel — GoldieBench AI benchmark screenshot▶ LIVE
Voxel 🥇
Visual
GLM built the densest, most detailed city — windowed skyscrapers, a speed + coins HUD. Opus ran the furthest with the cleanest motion (Score 303). Kimi's runner plays fine but is unforgiving — it crashes within seconds.
GLM-5.2 one-shot build of Waves — GoldieBench AI benchmark screenshot▶ LIVE
Waves
Visual
12KB · plays clean · three, webgl
every demo, in a grid · click any one to play

Compare GLM-5.2 against every other model

Every head-to-head featuring GLM-5.2. Verdicts shown for scored pairs.

GLM-5.2 vs Fusion
Fusion leads 39–5
GLM-5.2 vs Claude Opus 5
Claude Opus 5 leads 36–10
GLM-5.2 vs Hermes MoA
Hermes MoA leads 34–13
GLM-5.2 vs GPT-5.6 Sol
GPT-5.6 Sol leads 38–9
GLM-5.2 vs Claude Fable 5
Claude Fable 5 leads 32–12
GLM-5.2 vs Qwen 3.8
Qwen 3.8 leads 30–12
GLM-5.2 vs Grok
Grok leads 24–9
GLM-5.2 vs MiniMax M3
MiniMax M3 leads 28–11
GLM-5.2 vs Fugu Ultra
Fugu Ultra leads 23–15
GLM-5.2 vs Kimi K3
Kimi K3 leads 30–14
GLM-5.2 vs Fugu Mini
Fugu Mini leads 17–13
GLM-5.2 vs Opus 4.8
GLM-5.2 leads 20–9
GLM-5.2 vs Kimi K2.7
GLM-5.2 leads 10–6
GLM-5.2 vs Qwable 5 27B Coder
GLM-5.2 leads 27–14
GLM-5.2 vs Gemini 3.6 Flash
GLM-5.2 leads 27–20
GLM-5.2 vs Claude Sonnet 5
GLM-5.2 leads 23–22
GLM-5.2 vs Qwen 3.7
GLM-5.2 leads 29–6
GLM-5.2 vs Fugu Ultra 1.1
GLM-5.2 leads 12–11
GLM-5.2 vs Inkling
GLM-5.2 leads 43–4
GLM-5.2 vs Agents-A1
GLM-5.2 leads 41–1
GLM-5.2 vs Gemma 4 12B · MLX
GLM-5.2 leads 41–1
GLM-5.2 vs Laguna XS 2.1
GLM-5.2 leads 42–0
GLM-5.2 vs Qwythos 9B
GLM-5.2 leads 42–0
GLM-5.2 vs LongCat-2.0
LongCat-2.0 leads 2–1
GLM-5.2 vs Hy3
GLM-5.2 leads 7–0
GLM-5.2 vs Gemma-4 12B Coder
GLM-5.2 leads 6–0
GLM-5.2 vs DeepSeek V4 Flash
47 shared tasks · unscored
GLM-5.2 vs DeepSeek V4 Pro
47 shared tasks · unscored
GLM-5.2 vs Kimi K2.7 · Fast
47 shared tasks · unscored
GLM-5.2 vs Kimi K2.7 · No-Think
47 shared tasks · unscored
GLM-5.2 vs Kimi K2.7 · Quality
47 shared tasks · unscored
GLM-5.2 vs Ornith 1.0
42 shared tasks · unscored
GLM-5.2 vs Claude Mythos 5
Reference-only
GLM-5.2 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

GLM-5.2 vs Fusion GLM-5.2 vs Claude Opus 5 GLM-5.2 vs Hermes MoA GLM-5.2 vs GPT-5.6 Sol GLM-5.2 vs Claude Fable 5 GLM-5.2 vs Qwen 3.8 GLM-5.2 vs Grok GLM-5.2 vs MiniMax M3 GLM-5.2 vs Fugu Ultra GLM-5.2 vs Kimi K3 GLM-5.2 vs Fugu Mini GLM-5.2 vs Opus 4.8 GLM-5.2 vs Kimi K2.7 GLM-5.2 vs Qwable 5 27B Coder GLM-5.2 vs Gemini 3.6 Flash GLM-5.2 vs Claude Sonnet 5 GLM-5.2 vs Qwen 3.7 GLM-5.2 vs Fugu Ultra 1.1 GLM-5.2 vs Inkling GLM-5.2 vs Agents-A1 GLM-5.2 vs Gemma 4 12B · MLX GLM-5.2 vs Laguna XS 2.1 GLM-5.2 vs Qwythos 9B GLM-5.2 vs LongCat-2.0 GLM-5.2 vs Hy3 GLM-5.2 vs Gemma-4 12B Coder

Read more on agentos.guide: /glm-5-2, /glm-5-2-hermes, /glm-5-2-free-genius, /glm-5-2-benchmarks, /glm-vs-kimi-vs-opus, /glm-vs-qwen-vs-opus, /three-dragons, /the-content-machine, /the-everywhere-engine

GLM-5.2 — frequently asked

What is GLM-5.2?

GLM-5.2 is Zhipu / Z.ai's AI model — The never-forgets agent — 1M context, open weights. It has a 1M tokens context window and was released 2026-06-14.

How good is GLM-5.2 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 7.77/10 across 47 scored tasks, with 5 gold, 0 silver and 0 bronze medals.

How much does GLM-5.2 cost?

Open weights · free for individuals. Open-weights release: weights downloadable from Hugging Face for self-hosting, or runnable for free on z.ai for individuals (commercial use has separate licensing).

Where can I see GLM-5.2 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly