GoldieBench deep dive
Claude Opus 5's full benchmark breakdown.
The new Anthropic flagship — benched on all 45 one-shot builds the day it landed. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 13 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Gorgeous detailed player jet, full flight HUD, functional radar showing 4 bogeys, terrain warning, and afterburner/boost systems all read as a shippable 3D dogfighter. Slightly held back because the screenshot shows no visible enemy aircraft in view or active combat/tracers, leaving the on-screen action a touch quiet versus the field's best.”
“Gorgeous cohesive art direction — detailed low-poly car with wheels/shadow, layered cityscape, refined HUD with integrity/velocity/wanted stars and a live minimap showing cop (red) and player markers plus traffic (purple car visible). Pursuit=2 and full control set (fire/steal/boost) signal working combat, but no cops or firing visible in the frame keeps it just short of the very top.”
“Gorgeous atmospheric frozen world — aurora, moon glow, snowy terrain, low-poly pines, ruins, and a well-modeled armored third-person character with sword DRAWN and full HUD (compass, health/stam, kill counter). Strong and clearly shippable, but no visible enemies/dragon or combat action on screen keeps it just below the field's best; feels like a polished walking sim awaiting foes.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Claude Opus 5's lowest scored builds, verdicts unedited:
“The screenshot shows only raw JavaScript source text on a white page — nothing rendered, no canvas, no neon scene. Despite ambitious combat/vapor code in the source, the actual output is non-rendering.”
“The 3D scene fails to render — screen is blank except HUD, still showing 'KINDLING THE TORCH…' boot text. The source contains syntax errors (e.g. `0xffa martial=0`, `0x built=0`) that break the script, so no dungeon, torches, or enemies ever appear.”
“The HUD (health/stamina bars, ammo, wanted stars, minimap, controls) renders but the entire 3D scene is a blank flat purple void — no city, no hero, no pedestrians, no enemies visible, indicating the WebGL render failed or nothing is drawn. Despite ambitious source code, the actual screenshot shows a non-rendering game world.”
“The HUD renders beautifully (integrity/thrust/shield meters, score panel, crosshair, control hints), but the 3D scene is completely blank — no ground, mech, enemies, or landmarks visible, indicating the WebGL/three.js world failed to render. A source-code error (invalid `emissive:0x insertion` hex value in the material kit) likely broke script execution, leaving only the CSS HUD.”
05 · Category breakdown vs the whole field
Games (13 tasks)
Others (0 tasks)
Pages (0 tasks)
Sims (0 tasks)
Visuals (0 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Claude Opus 5 perform on GoldieBench?
Claude Opus 5 averages 5.60/10 across 13 scored one-shot build tasks, ranking #20 of 20 ranked frontier models, with 0 task golds, 0 silvers and 1 bronzes. Every score is a real judged build you can open and play on this site.
02What is Claude Opus 5 best at?
Its strongest scored build is Dogfight at 8.4/10. By category it averages Game 5.6 on the bench.
03How much does Claude Opus 5 cost?
Anthropic's brand-new flagship — the first Opus of the Claude 5 family, with a 1M-token context window. Benched via API the day it dropped; game tasks use our skill-infused AAA build prompts.
Source ledger
- 01Official vendor site: anthropic.com/claudewww.anthropic.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.