GoldieBench deep dive
Claude Opus 5's full benchmark breakdown.
The new Anthropic flagship — benched on all 45 one-shot builds the day it landed. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Stunning Schwarzschild ray-traced lensing with a convincing event horizon, warped accretion disk arcing over/under the shadow, Doppler beaming, and a rich lensed starfield — plus a polished, functional control panel with quality/inclination/distance sliders. The only faint quibble is a slightly hazy background and 27fps, but visually it matches the top of the field.”
“Stunning multi-type bursts (rings, peonies, chrysanthemums, willows) over a beautifully lit skyline with reflections, moon, and stars; rich UI with auto/finale/sound controls and live stats make it a clear task winner. Only minor concern is the 22fps under heavy load, but visually and functionally it's top-tier.”
“Gorgeous 3D N-body sim with a glowing central star, 167 orbiting bodies with velocity trails, merges counter (12), tilted disks and companion cluster — genuine physics with softening, substeps, and merging, plus a polished camera/UI. Strong: visible dynamic simulation and beautiful rendering; minor weak: 25fps at high body count.”
“Stunning multi-arm spiral with a glowing core, gradient star coloring (gold→pink→blue), halo, and distant starfield — visually the strongest galaxy render, backed by rich swirl/orbit/zoom/supernova interaction. Only nit is FPS at 39, but the composition clearly matches or beats the field's best.”
“Stunning depth with concentric glowing rings receding to a bright core, layered wireframe mesh, colorful motes, and radiating light streaks — genuinely reads as a 3D wormhole flythrough with polished HUD; steering/warp/regenerate controls back it up. Top of the field.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Claude Opus 5's lowest scored builds, verdicts unedited:
“Visually rich pool hall with detailed table, pendant lamps, spectator stools, and a working minimap plus combat feedback ("WRAITH CRACKED +120", enemies visible), but this is fundamentally an arena shooter reskin — it does NOT deliver a physically simulated billiards game as the brief demands, so it misses the core task despite strong 3D craft.”
“Visually impressive 3D neon arena with polished HUD, minimap, health/boost meters, combat feedback ('HULL HIT -9'), and hostiles on radar — but the brief explicitly asked for a CLASSIC arcade game (tetris/breakout/snake), and this is a 3D twin-stick shooter that ignores the pick entirely. Strong execution but off-brief.”
“Strong 3D presentation with polished HUD, minimap, ship model and space environment, but the brief asked for a 2D-style JUICY ARCADE space shooter with waves/bosses/power-ups — the screenshot shows a slow 3D cockpit-ish flight sim with faint asteroid-like triangles and no visible combat, projectiles, or enemy engagement (velocity 7 u/s, 0 kills). Reads more as an atmospheric flight demo than a punchy arcade blaster.”
“Strong polished HUD, textured raycaster walls, weapon viewmodel, and a detailed minimap show real engine work, plus counters imply combat (01/10 kills, 09 demons left). But the screenshot shows NO visible enemies on screen — combat can't be confirmed in the frame, so it reads closer to a walking sim than the enemy-chase brief demands.”
05 · Category breakdown vs the whole field
Games (23 tasks)
Others (3 tasks)
Pages (3 tasks)
Sims (12 tasks)
Visuals (9 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Claude Opus 5 perform on GoldieBench?
Claude Opus 5 averages 8.27/10 across 50 scored one-shot build tasks, ranking #2 of 23 ranked frontier models, with 13 task golds, 7 silvers and 5 bronzes. Every score is a real judged build you can open and play on this site.
02What is Claude Opus 5 best at?
Its strongest scored build is Blackhole at 9.0/10. By category it averages Game 8.0, Other 8.0, Page 8.5, Sim 8.6, Visual 8.5 on the bench.
03How much does Claude Opus 5 cost?
Anthropic's brand-new flagship — the first Opus of the Claude 5 family, with a 1M-token context window. Benched via API the day it dropped; game tasks use our skill-infused AAA build prompts.
Source ledger
- 01Official vendor site: anthropic.com/claudewww.anthropic.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.