GoldieBench deep dive
Claude Opus 5.5's full benchmark breakdown.
Anthropic's Opus 5.5 — benched on all 50 one-shot builds, every game playtested. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Striking real ray-marched Schwarzschild lensing: the Einstein ring bends over the top, the photon ring is thin and bright, the disk is Doppler-beamed and the streaked starfield looks good; controls are clean. Right up with the field's best, though the disk's banded ripple texture looks slightly synthetic.”
“Polished NovaOS desktop: working Paint, Notes with a sidebar and autosave, a Terminal with neofetch/ls, a menu bar with a live clock and CPU, a side launcher and a dock with running dots, all in one cohesive macOS-like style. Very strong, but not clearly above the field's 9.0 best.”
“Nails the look: a striped setting sun, neon wireframe mountains, palm-lined highway, a car with neon tail-lights, a chrome title and synth music on M. It is close to the field best but not clearly above it. The car is a dark low-detail silhouette.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are Claude Opus 5.5's lowest scored builds, verdicts unedited:
“EYEBALL GATE: the minimap sits over the title and HP/stamina/XP labels. The camera clips into a house wall, the scene is washed-out lavender-white, and the boxy wolf reads crude, though there is a quest tracker, a robed hero and rain.”
“EYEBALL GATE: the circular minimap covers the GOLD/KILLS/FOES/WAVE panel, so the HUD is not one cohesive layer. Heavy white bloom washes out the combat zone and a pillar hides the player, even though the hotbar, low-poly world and damage numbers are solid.”
“Hotbar, hearts, clock and a night-mob system exist, but the mid-play frame is crude: big flat-coloured blocks with no visible texture, the camera jammed against a tree trunk and leaves, and a cramped, hard-to-read view. Also carries the known pre-dusk burning-zombie glow bug.”
“EYEBALL GATE: the radar disc is drawn over the NEON BLASTER logo and the HULL/DASH labels in the top-left HUD. The crystal-asteroid set and ship model are decent, but the flat lavender floor has little of the neon glow or juice the brief asks for.”
05 · Category breakdown vs the whole field
Games (23 tasks)
Others (3 tasks)
Pages (3 tasks)
Sims (12 tasks)
Visuals (9 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
No externally sourced rows tracked for this model yet - the GoldieBench scores above are the record.
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does Claude Opus 5.5 perform on GoldieBench?
Claude Opus 5.5 averages 7.57/10 across 50 scored one-shot build tasks, ranking #14 of 26 ranked frontier models, with 0 task golds, 0 silvers and 1 bronzes. Every score is a real judged build you can open and play on this site.
02What is Claude Opus 5.5 best at?
Its strongest scored build is Blackhole at 8.8/10. By category it averages Game 7.1, Other 7.1, Page 7.8, Sim 8.1, Visual 8.0 on the bench.
03How much does Claude Opus 5.5 cost?
Anthropic's Opus 5.5 — 1M-token context, listed at $4 input / $20 output per million tokens. Benched through a Claude subscription; game tasks use our skill-infused AAA build prompts.
Source ledger
- 01Official vendor site: anthropic.com/claudewww.anthropic.com
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.