Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Anthropic

Opus 4.8

The reasoning king — deepest thinking, premium price.

Context200,000 tokens (1M with extended thinking)
Pricing$15 / $75 per M tokens
Tasks tested47
Avg score7.51/10 average
Medals🥇3 🥈1 🥉1
Release2026-05
Official vendor source
Opus 4.8 is built by Anthropic — see the vendor's own product page, pricing, and docs at anthropic.com/claude.
Visit anthropic.com/claude →

Reference benchmarks for Opus 4.8

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Opus 4.8 is honest about what's measured.

SWE-bench Verified
88.6%

What is Opus 4.8?

Opus 4.8 is the Anthropic frontier model with a 200,000 tokens (1M with extended thinking) context window, released 2026-05. Tagline: The reasoning king — deepest thinking, premium price.. Official source: anthropic.com/claude.

Pricing detail. Premium pricing via the Anthropic API: $15 per million input tokens, $75 per million output tokens. Extended thinking is included but adds latency.

How I use it inside the Agent OS. The default when the build has to ship on the first prompt — Opus is the safety net inside Agent OS for hard one-shots.

What I built with Opus 4.8

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Opus 4.8 shipped on the bench: 47 one-shot demos across 200,000 tokens (1M with extended thinking) of context. Of those, 47 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Most consistent across the Goldie Bench bench — no weak build, 8.46/10 average
  • Deepest one-shot reasoning, especially on game-feel and physics
  • Extended thinking mode handles up to 1M tokens of context

Trade-offs

  • 5–10× the per-token cost of every other model on the bench
  • Less flair on cinematic visuals than GLM-5.2 — playing it safer wins on accuracy, costs you on showpiece moments

Best for

  • Mission-critical one-shot builds where 'has to work the first time' matters
  • Hard reasoning tasks (planning, multi-step) where you'll pay for the depth
  • Anything where vendor reliability beats the per-token bill

Every demo by Opus 4.8

47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Opus 4.8 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
13KB · animation runs but no input response · webgl
Opus 4.8 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
13KB · plays clean · webgl
Opus 4.8 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom 🥈
Game
All three are real, playable shooters. Opus drops you in a corridor with an imp dead ahead — gun, crosshair and HUD framed like a screenshot. Kimi matches it: a monster down a textured hall, health, ammo, minimap. GLM ships a gorgeous 'HAZARD PROTOCOL' title screen with a working game behind it, though it too spawns facing a wall. Opus by a hair on the cleanest fight.
Opus 4.8 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
16KB · plays clean · webgl
Opus 4.8 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
11KB · plays clean · webgl
Opus 4.8 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
22KB · plays clean · webgl
Opus 4.8 one-shot build of Voxelcraft — GoldieBench AI benchmark screenshot▶ LIVE
Voxelcraft
Game
7KB · plays clean · webgl, input, rAF
Opus 4.8 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
20KB · plays clean · three, webgl (re-rolled)
Opus 4.8 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
22KB · plays clean · three, webgl (re-rolled)
Opus 4.8 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo
Page
14KB · plays clean · plain
Opus 4.8 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute
Game
21KB · plays clean · three, webgl
Opus 4.8 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim
Game
27KB · plays clean · three, webgl
Opus 4.8 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade 🥉
Game
All three shipped a genuinely juicy game. Opus's breakout had the most game-feel — particle bursts and a live combo. Kimi's breakout was clean and solid. GLM went its own way with fullscreen neon asteroids. The closest of the practical five.
Opus 4.8 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
14KB · plays clean · webgl, input
Opus 4.8 one-shot build of Dragonflight — GoldieBench AI benchmark screenshot▶ LIVE
Dragonflight
Game
12KB · plays clean · webgl, input
Opus 4.8 one-shot build of Game — GoldieBench AI benchmark screenshot▶ LIVE
Game
Game
17KB · plays clean · audio
Opus 4.8 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
16KB · plays clean · audio, input
Opus 4.8 one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity
Game
GLM's is the most cinematic — neon towers, a setting sun, Japanese signage and a flight HUD, like a frame from a film. Opus's is a clean canyon of lit skyscrapers racing to a vanishing point. Kimi leaned into the synthwave sun and grid more than the city itself. GLM wins the skyline.
Opus 4.8 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
9KB · plays clean · input
Opus 4.8 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
14KB · animation runs but no input response · webgl, bloom
Opus 4.8 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
GLM shipped the full arcade package — an 'OUTRUN 2086' title, gear, RPM and velocity dials, mountains, the car cruising at 90+. Opus's road curves hard past rumble strips and palms into a scanline sun. Kimi's 'NEON OUTRUN' is clean and on-brief. GLM edges it on sheer completeness.
Opus 4.8 one-shot build of Pool — GoldieBench AI benchmark screenshot▶ LIVE
Pool
Game
9KB · animation runs but no input response · plain
Opus 4.8 one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster
Game
Kimi nailed it — brick walls, a checkered floor, a clean minimap, textbook Wolfenstein, runs clean out of the box. Opus's is close and more atmospheric: warm fog and a vignette down a stone corridor (A/D to turn, W/S to move). GLM's engine is genuinely good — brick and mossy-stone walls, fog, a minimap — but its one-shot spawned the player buried inside a wall, dead on arrival; I nudged the start one cell so you can actually walk it. That spawn bug is why it scores lowest here, even though the e
Opus 4.8 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
11KB · plays clean · input
Opus 4.8 one-shot build of Landing — GoldieBench AI benchmark screenshot▶ LIVE
Landing 🥇
Page
Funniest result of the lot: GLM and Opus independently produced near-identical premium 'Introducing Nova 1 — Intelligence, reimagined / distilled' keynote pages — gradient hero, full nav, pricing tiers. A dead heat. Kimi's was a plainer set of feature cards.
Opus 4.8 one-shot build of Webos — GoldieBench AI benchmark screenshot▶ LIVE
Webos
Page
16KB · animation runs but no input response · plain
Opus 4.8 one-shot build of Blackhole — GoldieBench AI benchmark screenshot▶ LIVE
Blackhole 🥇
Sim
Opus nailed it — a pure-black event horizon, a bright photon ring, and the disk bent up and over the top exactly like the film's lensing. GLM came in strong with a clean ring and a starfield warping past the hole. Kimi's disk is fine, but the background is a soft grey blur instead of stars. This one's Opus's.
Opus 4.8 one-shot build of Boids — GoldieBench AI benchmark screenshot▶ LIVE
Boids
Sim
8KB · plays clean · plain
Opus 4.8 one-shot build of Cloth — GoldieBench AI benchmark screenshot▶ LIVE
Cloth
Sim
3KB · plays clean · rAF
Opus 4.8 one-shot build of Fluid — GoldieBench AI benchmark screenshot▶ LIVE
Fluid
Sim
GLM filled the bowl with glowing liquid that actually sloshes — the most convincing 'liquid in a bowl'. Opus's particles glowed but clumped to the centre. Kimi's collapsed into a tiny blob.
Opus 4.8 one-shot build of Fractal — GoldieBench AI benchmark screenshot▶ LIVE
Fractal
Sim
All three are genuinely good. Kimi's is the jaw-dropper — a deep rainbow plunge into a seahorse spiral, dense with self-similar detail. Opus zooms smoothly into the seahorse valley with a tasteful cycling palette. GLM frames the whole iconic set in a fire palette with a live coordinate HUD, then descends. Kimi takes this one on raw spectacle.
Opus 4.8 one-shot build of Galaxy — GoldieBench AI benchmark screenshot▶ LIVE
Galaxy
Sim
Opus built a proper interactive 3D galaxy — drag to orbit a 7,000-star cloud around a glowing core. Kimi's is the prettiest single frame: a clean tilted spiral disk with rainbow arms. GLM's runs on a canvas with a slick NGC-style HUD and zoom, just less dramatic at a glance. Three good galaxies, three different bets.
Opus 4.8 one-shot build of Orbit — GoldieBench AI benchmark screenshot▶ LIVE
Orbit 🥇
Sim
Opus nailed the brief — labelled planet orbits, a real NEO / close-pass panel, a sim clock. GLM went for drama: a glowing nebula swirl that's gorgeous but reads more galaxy than orbit map. Kimi's is accurate but dim and sparse.
Opus 4.8 one-shot build of Particleforge — GoldieBench AI benchmark screenshot▶ LIVE
Particleforge
Sim
10KB · plays clean · plain
Opus 4.8 one-shot build of Pathtracer — GoldieBench AI benchmark screenshot▶ LIVE
Pathtracer
Sim
6KB · plays clean · webgl, rAF
Opus 4.8 one-shot build of Reactiondiff — GoldieBench AI benchmark screenshot▶ LIVE
Reactiondiff
Sim
5KB · plays clean · webgl, rAF
Opus 4.8 one-shot build of Solar — GoldieBench AI benchmark screenshot▶ LIVE
Solar
Sim
Three genuinely good space sims. Opus tilts the orbits into real 3D with a bloom-heavy sun and Saturn's rings. GLM's is the most product-like — labelled planets, orbit and label toggles, a clean HUD. Kimi's is a tidy tilted-orbit system with rings and a deep starfield. Opus and GLM are neck-and-neck; Opus takes it on the 3D feel.
Opus 4.8 one-shot build of Wormhole — GoldieBench AI benchmark screenshot▶ LIVE
Wormhole
Sim
11KB · plays clean · webgl, input
Opus 4.8 one-shot build of Aurora — GoldieBench AI benchmark screenshot▶ LIVE
Aurora
Visual
4KB · animation runs but no input response · rAF
Opus 4.8 one-shot build of Fireworks — GoldieBench AI benchmark screenshot▶ LIVE
Fireworks
Visual
5KB · plays clean · rAF
Opus 4.8 one-shot build of Lavalamp — GoldieBench AI benchmark screenshot▶ LIVE
Lavalamp
Visual
3KB · plays clean · webgl, rAF
Opus 4.8 one-shot build of Matrix — GoldieBench AI benchmark screenshot▶ LIVE
Matrix
Visual
2KB · plays clean · rAF
Opus 4.8 one-shot build of Plasma — GoldieBench AI benchmark screenshot▶ LIVE
Plasma
Visual
5KB · plays clean · webgl, rAF
Opus 4.8 one-shot build of Synthwave — GoldieBench AI benchmark screenshot▶ LIVE
Synthwave
Visual
This is GLM's. A cyan wireframe mountain range scrolling under a scanline synthwave sun — the single most beautiful frame in the whole shoot-out. Opus's clean Tron grid and magenta horizon is a close, cooler-toned second. Kimi got the idea but blew the exposure — the grid washes out to near-white. GLM wins this one going away.
Opus 4.8 one-shot build of Terrain — GoldieBench AI benchmark screenshot▶ LIVE
Terrain
Visual
2KB · plays clean · rAF
Opus 4.8 one-shot build of Voxel — GoldieBench AI benchmark screenshot▶ LIVE
Voxel
Visual
GLM built the densest, most detailed city — windowed skyscrapers, a speed + coins HUD. Opus ran the furthest with the cleanest motion (Score 303). Kimi's runner plays fine but is unforgiving — it crashes within seconds.
Opus 4.8 one-shot build of Waves — GoldieBench AI benchmark screenshot▶ LIVE
Waves
Visual
6KB · plays clean · three, webgl, controls, rAF
every demo, in a grid · click any one to play

Compare Opus 4.8 against every other model

Every head-to-head featuring Opus 4.8. Verdicts shown for scored pairs.

Opus 4.8 vs Fusion
Fusion leads 36–2
Opus 4.8 vs Claude Opus 5
Claude Opus 5 leads 36–9
Opus 4.8 vs Hermes MoA
Hermes MoA leads 35–11
Opus 4.8 vs GPT-5.6 Sol
GPT-5.6 Sol leads 40–7
Opus 4.8 vs Claude Fable 5
Claude Fable 5 leads 28–14
Opus 4.8 vs Qwen 3.8
Qwen 3.8 leads 33–9
Opus 4.8 vs Grok
Grok leads 22–8
Opus 4.8 vs MiniMax M3
MiniMax M3 leads 27–14
Opus 4.8 vs Fugu Ultra
Fugu Ultra leads 25–11
Opus 4.8 vs Kimi K3
Kimi K3 leads 33–13
Opus 4.8 vs GLM-5.2
GLM-5.2 leads 20–9
Opus 4.8 vs Fugu Mini
Fugu Mini leads 20–10
Opus 4.8 vs Kimi K2.7
Opus 4.8 leads 12–6
Opus 4.8 vs Qwable 5 27B Coder
Opus 4.8 leads 21–20
Opus 4.8 vs Gemini 3.6 Flash
Gemini 3.6 Flash leads 24–23
Opus 4.8 vs Claude Sonnet 5
Claude Sonnet 5 leads 26–20
Opus 4.8 vs Qwen 3.7
Opus 4.8 leads 21–11
Opus 4.8 vs Fugu Ultra 1.1
Opus 4.8 leads 12–11
Opus 4.8 vs Inkling
Opus 4.8 leads 41–6
Opus 4.8 vs Agents-A1
Opus 4.8 leads 34–8
Opus 4.8 vs Gemma 4 12B · MLX
Opus 4.8 leads 38–3
Opus 4.8 vs Laguna XS 2.1
Opus 4.8 leads 41–1
Opus 4.8 vs Qwythos 9B
Opus 4.8 leads 41–1
Opus 4.8 vs LongCat-2.0
LongCat-2.0 leads 3–1
Opus 4.8 vs Hy3
Opus 4.8 leads 5–2
Opus 4.8 vs Gemma-4 12B Coder
Opus 4.8 leads 6–0
Opus 4.8 vs DeepSeek V4 Flash
47 shared tasks · unscored
Opus 4.8 vs DeepSeek V4 Pro
47 shared tasks · unscored
Opus 4.8 vs Kimi K2.7 · Fast
47 shared tasks · unscored
Opus 4.8 vs Kimi K2.7 · No-Think
47 shared tasks · unscored
Opus 4.8 vs Kimi K2.7 · Quality
47 shared tasks · unscored
Opus 4.8 vs Ornith 1.0
42 shared tasks · unscored
Opus 4.8 vs Claude Mythos 5
Reference-only
Opus 4.8 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Opus 4.8 vs Fusion Opus 4.8 vs Claude Opus 5 Opus 4.8 vs Hermes MoA Opus 4.8 vs GPT-5.6 Sol Opus 4.8 vs Claude Fable 5 Opus 4.8 vs Qwen 3.8 Opus 4.8 vs Grok Opus 4.8 vs MiniMax M3 Opus 4.8 vs Fugu Ultra Opus 4.8 vs Kimi K3 Opus 4.8 vs GLM-5.2 Opus 4.8 vs Fugu Mini Opus 4.8 vs Kimi K2.7 Opus 4.8 vs Qwable 5 27B Coder Opus 4.8 vs Gemini 3.6 Flash Opus 4.8 vs Claude Sonnet 5 Opus 4.8 vs Qwen 3.7 Opus 4.8 vs Fugu Ultra 1.1 Opus 4.8 vs Inkling Opus 4.8 vs Agents-A1 Opus 4.8 vs Gemma 4 12B · MLX Opus 4.8 vs Laguna XS 2.1 Opus 4.8 vs Qwythos 9B Opus 4.8 vs LongCat-2.0 Opus 4.8 vs Hy3 Opus 4.8 vs Gemma-4 12B Coder

Read more on agentos.guide: /opus-ultracode, /claude-fable-5, /glm-vs-kimi-vs-opus, /glm-vs-qwen-vs-opus

Opus 4.8 — frequently asked

What is Opus 4.8?

Opus 4.8 is Anthropic's AI model — The reasoning king — deepest thinking, premium price. It has a 200K tokens context window and was released 2026-05.

How good is Opus 4.8 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 7.51/10 across 47 scored tasks, with 3 gold, 1 silver and 1 bronze medals.

How much does Opus 4.8 cost?

$15 / $75 per M tokens. Premium pricing via the Anthropic API: $15 per million input tokens, $75 per million output tokens. Extended thinking is included but adds latency.

Where can I see Opus 4.8 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly