Claude Opus 5.5
Anthropic's Opus 5.5 — benched on all 50 one-shot builds, every game playtested.
What is Claude Opus 5.5?
Claude Opus 5.5 is the Anthropic frontier model with a 1,000,000 tokens context window, released 2026-09. Tagline: Anthropic's Opus 5.5 — benched on all 50 one-shot builds, every game playtested.. Official source: anthropic.com/claude.
Pricing detail. Anthropic's Opus 5.5 — 1M-token context, listed at $4 input / $20 output per million tokens. Benched through a Claude subscription; game tasks use our skill-infused AAA build prompts.
How I use it inside the Agent OS. Benched on all 50 GoldieBench tasks via a Claude subscription: one-shot builds, real renders, pixel-diff playtests on every game, then vision-judged.
What I built with Claude Opus 5.5
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Claude Opus 5.5 shipped on the bench: 50 one-shot demos across 1,000,000 tokens of context. Of those, 50 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Frontier-class coding + agentic reasoning (Claude 5 family)
- 1M-token context — reads an entire codebase in one call
- All 50 one-shot builds rendered with zero console errors; all 23 games passed the input playtest
Trade-offs
- Premium pricing ($4/$20 per M) — route the everyday 90% to cheaper lanes
- One-shot builds can still ship logic bugs (e.g. a doom kill counter that miscounts)
Best for
- Hardest agentic builds
- Whole-repo reasoning
- Frontier one-shots
Every benchmark — Claude Opus 5.5's full scorecard
All 50 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Claude Opus 5.5 deep dive →.
Every demo by Claude Opus 5.5
50 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Claude Opus 5.5 against every other model
Every head-to-head featuring Claude Opus 5.5. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Claude Opus 5.5 vs Fusion Claude Opus 5.5 vs Claude Opus 5 Claude Opus 5.5 vs Hermes MoA Claude Opus 5.5 vs GPT-5.6 Sol Claude Opus 5.5 vs Claude Fable 5 Claude Opus 5.5 vs Qwen 3.8 Claude Opus 5.5 vs Grok Claude Opus 5.5 vs MiniMax M3 Claude Opus 5.5 vs Fugu Ultra Claude Opus 5.5 vs Kimi K3 Claude Opus 5.5 vs GLM-5.2 Claude Opus 5.5 vs Fugu Mini Claude Opus 5.5 vs Muse Spark 1.2 Claude Opus 5.5 vs Opus 4.8 Claude Opus 5.5 vs Kimi K2.7 Claude Opus 5.5 vs Grok 4.7 Claude Opus 5.5 vs Qwable 5 27B Coder Claude Opus 5.5 vs Gemini 3.6 Flash Claude Opus 5.5 vs Claude Sonnet 5 Claude Opus 5.5 vs Qwen 3.7 Claude Opus 5.5 vs Fugu Ultra 1.1 Claude Opus 5.5 vs MiMo-V2.6 Pro Claude Opus 5.5 vs Inkling Claude Opus 5.5 vs Grok 4.6 Claude Opus 5.5 vs Agents-A1 Claude Opus 5.5 vs Gemma 4 12B · MLX Claude Opus 5.5 vs Laguna XS 2.1 Claude Opus 5.5 vs Qwythos 9B Claude Opus 5.5 vs LongCat-2.0 Claude Opus 5.5 vs Hy3 Claude Opus 5.5 vs Gemma-4 12B CoderRead more on agentos.guide:
Claude Opus 5.5 — frequently asked
What is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's AI model — Anthropic's Opus 5.5 — benched on all 50 one-shot builds, every game playtested. It has a 1M tokens context window and was released 2026-09.
How good is Claude Opus 5.5 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 7.57/10 across 50 scored tasks, with 0 gold, 0 silver and 1 bronze medals.
How much does Claude Opus 5.5 cost?
$4 / $20 per M. Anthropic's Opus 5.5 — 1M-token context, listed at $4 input / $20 output per million tokens. Benched through a Claude subscription; game tasks use our skill-infused AAA build prompts.
Where can I see Claude Opus 5.5 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.