Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Claude Opus 5.5's full benchmark breakdown.

Anthropic's Opus 5.5 — benched on all 50 one-shot builds, every game playtested. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-09-23
Scored tasks
50
Reading time
10 min
External sources
1

01 · The headline numbers

GoldieBench average
7.57/10
50 scored one-shot tasks
Board rank
#14
of 26 ranked frontier models
Task medals
0🥇 0🥈 1🥉
outright wins on shared briefs
Context
1,000,000 tokens
$4 / $20 per M

02 · Every benchmark, bar by bar

All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Striking real ray-marched Schwarzschild lensing: the Einstein ring bends over the top, the photon ring is thin and bright, the disk is Doppler-beamed and the streaked starfield looks good; controls are clean. Right up with the field's best, though the disk's banded ripple texture looks slightly synthetic.”

- the judge's verdict, unedited

“Polished NovaOS desktop: working Paint, Notes with a sidebar and autosave, a Terminal with neofetch/ls, a menu bar with a live clock and CPU, a side launcher and a dock with running dots, all in one cohesive macOS-like style. Very strong, but not clearly above the field's 9.0 best.”

- the judge's verdict, unedited

“Nails the look: a striped setting sun, neon wireframe mountains, palm-lined highway, a car with neon tail-lights, a chrome title and synth music on M. It is close to the field best but not clearly above it. The car is a dark low-detail silhouette.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Claude Opus 5.5's lowest scored builds, verdicts unedited:

“EYEBALL GATE: the minimap sits over the title and HP/stamina/XP labels. The camera clips into a house wall, the scene is washed-out lavender-white, and the boxy wolf reads crude, though there is a quest tracker, a robed hero and rain.”

- the judge's verdict, unedited

“EYEBALL GATE: the circular minimap covers the GOLD/KILLS/FOES/WAVE panel, so the HUD is not one cohesive layer. Heavy white bloom washes out the combat zone and a pillar hides the player, even though the hotbar, low-poly world and damage numbers are solid.”

- the judge's verdict, unedited

“Hotbar, hearts, clock and a night-mob system exist, but the mid-play frame is crude: big flat-coloured blocks with no visible texture, the camera jammed against a tree trunk and leaves, and a cramped, hard-to-read view. Also carries the known pre-dusk burning-zombie glow bug.”

- the judge's verdict, unedited

“EYEBALL GATE: the radar disc is drawn over the NEON BLASTER logo and the HULL/DASH labels in the top-left HUD. The crystal-asteroid set and ship model are decent, but the flat lavender floor has little of the neon glow or juice the brief asks for.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (23 tasks)

Claude Opus 5.57.1
field avg7.1

Others (3 tasks)

Claude Opus 5.57.1
field avg7.4

Pages (3 tasks)

Claude Opus 5.57.8
field avg7.7

Sims (12 tasks)

Claude Opus 5.58.1
field avg6.9

Visuals (9 tasks)

Claude Opus 5.58.0
field avg7.0

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Claude Opus 5.5 perform on GoldieBench?

Claude Opus 5.5 averages 7.57/10 across 50 scored one-shot build tasks, ranking #14 of 26 ranked frontier models, with 0 task golds, 0 silvers and 1 bronzes. Every score is a real judged build you can open and play on this site.

02What is Claude Opus 5.5 best at?

Its strongest scored build is Blackhole at 8.8/10. By category it averages Game 7.1, Other 7.1, Page 7.8, Sim 8.1, Visual 8.0 on the bench.

03How much does Claude Opus 5.5 cost?

Anthropic's Opus 5.5 — 1M-token context, listed at $4 input / $20 output per million tokens. Benched through a Claude subscription; game tasks use our skill-infused AAA build prompts.

Source ledger

  1. 01Official vendor site: anthropic.com/claudewww.anthropic.com

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly