GLM-5.2
The never-forgets agent — 1M context, open weights.
Reference benchmarks for GLM-5.2
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for GLM-5.2 is honest about what's measured.
What is GLM-5.2?
GLM-5.2 is the Zhipu / Z.ai frontier model with a 1,000,000 tokens context window, released 2026-06-14. Tagline: The never-forgets agent — 1M context, open weights.. Official source: z.ai.
Pricing detail. Open-weights release: weights downloadable from Hugging Face for self-hosting, or runnable for free on z.ai for individuals (commercial use has separate licensing).
How I use it inside the Agent OS. Default model inside Agent OS for any task that touches a long context — codebase Q&A, multi-file refactors, agent memory replay.
What I built with GLM-5.2
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what GLM-5.2 shipped on the bench: 47 one-shot demos across 1,000,000 tokens of context. Of those, 47 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- 1M-token context window — best-in-class long-document and large-codebase work
- Open weights — runs locally, no vendor lock-in, no token meter
- Top of the bench for cinematic visuals (neon city, synthwave, voxel runner)
Trade-offs
- Faceplanted on the Goldie Bench raycaster — the engine was great but it spawned the player inside a wall
- First-shot reliability lags Opus by a hair on consistency
Best for
- Long-context agent loops — pasting a whole codebase into one prompt
- Cinematic visual builds — landing pages, voxel scenes, synthwave runners
- Anyone who needs to run a frontier coder locally for $0
Every benchmark — GLM-5.2's full scorecard
All 47 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the GLM-5.2 deep dive →.
Every demo by GLM-5.2
47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare GLM-5.2 against every other model
Every head-to-head featuring GLM-5.2. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
GLM-5.2 vs Fusion GLM-5.2 vs Claude Opus 5 GLM-5.2 vs Hermes MoA GLM-5.2 vs GPT-5.6 Sol GLM-5.2 vs Claude Fable 5 GLM-5.2 vs Qwen 3.8 GLM-5.2 vs Grok GLM-5.2 vs MiniMax M3 GLM-5.2 vs Fugu Ultra GLM-5.2 vs Kimi K3 GLM-5.2 vs Fugu Mini GLM-5.2 vs Opus 4.8 GLM-5.2 vs Kimi K2.7 GLM-5.2 vs Qwable 5 27B Coder GLM-5.2 vs Gemini 3.6 Flash GLM-5.2 vs Claude Sonnet 5 GLM-5.2 vs Qwen 3.7 GLM-5.2 vs Fugu Ultra 1.1 GLM-5.2 vs Inkling GLM-5.2 vs Agents-A1 GLM-5.2 vs Gemma 4 12B · MLX GLM-5.2 vs Laguna XS 2.1 GLM-5.2 vs Qwythos 9B GLM-5.2 vs LongCat-2.0 GLM-5.2 vs Hy3 GLM-5.2 vs Gemma-4 12B CoderRead more on agentos.guide: /glm-5-2, /glm-5-2-hermes, /glm-5-2-free-genius, /glm-5-2-benchmarks, /glm-vs-kimi-vs-opus, /glm-vs-qwen-vs-opus, /three-dragons, /the-content-machine, /the-everywhere-engine
GLM-5.2 — frequently asked
What is GLM-5.2?
GLM-5.2 is Zhipu / Z.ai's AI model — The never-forgets agent — 1M context, open weights. It has a 1M tokens context window and was released 2026-06-14.
How good is GLM-5.2 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 7.77/10 across 47 scored tasks, with 5 gold, 0 silver and 0 bronze medals.
How much does GLM-5.2 cost?
Open weights · free for individuals. Open-weights release: weights downloadable from Hugging Face for self-hosting, or runnable for free on z.ai for individuals (commercial use has separate licensing).
Where can I see GLM-5.2 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.