Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Claude Opus 5's full benchmark breakdown.

The new Anthropic flagship — benched on all 45 one-shot builds the day it landed. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-08-13
Scored tasks
50
Reading time
10 min
External sources
1

01 · The headline numbers

GoldieBench average
8.27/10
50 scored one-shot tasks
Board rank
#2
of 23 ranked frontier models
Task medals
13🥇 7🥈 5🥉
outright wins on shared briefs
Context
1,000,000 tokens
$5 / $25 per M

02 · Every benchmark, bar by bar

All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Stunning Schwarzschild ray-traced lensing with a convincing event horizon, warped accretion disk arcing over/under the shadow, Doppler beaming, and a rich lensed starfield — plus a polished, functional control panel with quality/inclination/distance sliders. The only faint quibble is a slightly hazy background and 27fps, but visually it matches the top of the field.”

- the judge's verdict, unedited

“Stunning multi-type bursts (rings, peonies, chrysanthemums, willows) over a beautifully lit skyline with reflections, moon, and stars; rich UI with auto/finale/sound controls and live stats make it a clear task winner. Only minor concern is the 22fps under heavy load, but visually and functionally it's top-tier.”

- the judge's verdict, unedited

“Gorgeous 3D N-body sim with a glowing central star, 167 orbiting bodies with velocity trails, merges counter (12), tilted disks and companion cluster — genuine physics with softening, substeps, and merging, plus a polished camera/UI. Strong: visible dynamic simulation and beautiful rendering; minor weak: 25fps at high body count.”

- the judge's verdict, unedited

“Stunning multi-arm spiral with a glowing core, gradient star coloring (gold→pink→blue), halo, and distant starfield — visually the strongest galaxy render, backed by rich swirl/orbit/zoom/supernova interaction. Only nit is FPS at 39, but the composition clearly matches or beats the field's best.”

- the judge's verdict, unedited

“Stunning depth with concentric glowing rings receding to a bright core, layered wireframe mesh, colorful motes, and radiating light streaks — genuinely reads as a 3D wormhole flythrough with polished HUD; steering/warp/regenerate controls back it up. Top of the field.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Claude Opus 5's lowest scored builds, verdicts unedited:

“Visually rich pool hall with detailed table, pendant lamps, spectator stools, and a working minimap plus combat feedback ("WRAITH CRACKED +120", enemies visible), but this is fundamentally an arena shooter reskin — it does NOT deliver a physically simulated billiards game as the brief demands, so it misses the core task despite strong 3D craft.”

- the judge's verdict, unedited

“Visually impressive 3D neon arena with polished HUD, minimap, health/boost meters, combat feedback ('HULL HIT -9'), and hostiles on radar — but the brief explicitly asked for a CLASSIC arcade game (tetris/breakout/snake), and this is a 3D twin-stick shooter that ignores the pick entirely. Strong execution but off-brief.”

- the judge's verdict, unedited

“Strong 3D presentation with polished HUD, minimap, ship model and space environment, but the brief asked for a 2D-style JUICY ARCADE space shooter with waves/bosses/power-ups — the screenshot shows a slow 3D cockpit-ish flight sim with faint asteroid-like triangles and no visible combat, projectiles, or enemy engagement (velocity 7 u/s, 0 kills). Reads more as an atmospheric flight demo than a punchy arcade blaster.”

- the judge's verdict, unedited

“Strong polished HUD, textured raycaster walls, weapon viewmodel, and a detailed minimap show real engine work, plus counters imply combat (01/10 kills, 09 demons left). But the screenshot shows NO visible enemies on screen — combat can't be confirmed in the frame, so it reads closer to a walking sim than the enemy-chase brief demands.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (23 tasks)

Claude Opus 58.0
field avg7.1

Others (3 tasks)

Claude Opus 58.0
field avg7.5

Pages (3 tasks)

Claude Opus 58.5
field avg7.6

Sims (12 tasks)

Claude Opus 58.6
field avg6.9

Visuals (9 tasks)

Claude Opus 58.5
field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Claude Opus 5 perform on GoldieBench?

Claude Opus 5 averages 8.27/10 across 50 scored one-shot build tasks, ranking #2 of 23 ranked frontier models, with 13 task golds, 7 silvers and 5 bronzes. Every score is a real judged build you can open and play on this site.

02What is Claude Opus 5 best at?

Its strongest scored build is Blackhole at 9.0/10. By category it averages Game 8.0, Other 8.0, Page 8.5, Sim 8.6, Visual 8.5 on the bench.

03How much does Claude Opus 5 cost?

Anthropic's brand-new flagship — the first Opus of the Claude 5 family, with a 1M-token context window. Benched via API the day it dropped; game tasks use our skill-infused AAA build prompts.

Source ledger

  1. 01Official vendor site: anthropic.com/claudewww.anthropic.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly