Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Claude Opus 5's full benchmark breakdown.

The new Anthropic flagship — benched on all 45 one-shot builds the day it landed. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-07-24
Scored tasks
13
Reading time
5 min
External sources
1

01 · The headline numbers

GoldieBench average
5.60/10
13 scored one-shot tasks
Board rank
#20
of 20 ranked frontier models
Task medals
0🥇 0🥈 1🥉
outright wins on shared briefs
Context
1,000,000 tokens
$5 / $25 per M

02 · Every benchmark, bar by bar

All 13 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Gorgeous detailed player jet, full flight HUD, functional radar showing 4 bogeys, terrain warning, and afterburner/boost systems all read as a shippable 3D dogfighter. Slightly held back because the screenshot shows no visible enemy aircraft in view or active combat/tracers, leaving the on-screen action a touch quiet versus the field's best.”

- the judge's verdict, unedited

“Gorgeous cohesive art direction — detailed low-poly car with wheels/shadow, layered cityscape, refined HUD with integrity/velocity/wanted stars and a live minimap showing cop (red) and player markers plus traffic (purple car visible). Pursuit=2 and full control set (fire/steal/boost) signal working combat, but no cops or firing visible in the frame keeps it just short of the very top.”

- the judge's verdict, unedited

“Gorgeous atmospheric frozen world — aurora, moon glow, snowy terrain, low-poly pines, ruins, and a well-modeled armored third-person character with sword DRAWN and full HUD (compass, health/stam, kill counter). Strong and clearly shippable, but no visible enemies/dragon or combat action on screen keeps it just below the field's best; feels like a polished walking sim awaiting foes.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Claude Opus 5's lowest scored builds, verdicts unedited:

“The screenshot shows only raw JavaScript source text on a white page — nothing rendered, no canvas, no neon scene. Despite ambitious combat/vapor code in the source, the actual output is non-rendering.”

- the judge's verdict, unedited

“The 3D scene fails to render — screen is blank except HUD, still showing 'KINDLING THE TORCH…' boot text. The source contains syntax errors (e.g. `0xffa martial=0`, `0x built=0`) that break the script, so no dungeon, torches, or enemies ever appear.”

- the judge's verdict, unedited

“The HUD (health/stamina bars, ammo, wanted stars, minimap, controls) renders but the entire 3D scene is a blank flat purple void — no city, no hero, no pedestrians, no enemies visible, indicating the WebGL render failed or nothing is drawn. Despite ambitious source code, the actual screenshot shows a non-rendering game world.”

- the judge's verdict, unedited

“The HUD renders beautifully (integrity/thrust/shield meters, score panel, crosshair, control hints), but the 3D scene is completely blank — no ground, mech, enemies, or landmarks visible, indicating the WebGL/three.js world failed to render. A source-code error (invalid `emissive:0x insertion` hex value in the material kit) likely broke script execution, leaving only the CSS HUD.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (13 tasks)

Claude Opus 55.6
field avg7.1

Others (0 tasks)

field avg7.5

Pages (0 tasks)

field avg7.6

Sims (0 tasks)

field avg6.7

Visuals (0 tasks)

field avg6.8

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Claude Opus 5 perform on GoldieBench?

Claude Opus 5 averages 5.60/10 across 13 scored one-shot build tasks, ranking #20 of 20 ranked frontier models, with 0 task golds, 0 silvers and 1 bronzes. Every score is a real judged build you can open and play on this site.

02What is Claude Opus 5 best at?

Its strongest scored build is Dogfight at 8.4/10. By category it averages Game 5.6 on the bench.

03How much does Claude Opus 5 cost?

Anthropic's brand-new flagship — the first Opus of the Claude 5 family, with a 1M-token context window. Benched via API the day it dropped; game tasks use our skill-infused AAA build prompts.

Source ledger

  1. 01Official vendor site: anthropic.com/claudewww.anthropic.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly