Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Game

Game

Game — generic 'make a game' open prompt.

CategoryGame
Models tested24
Scored18/24
Avg score8.11/10
WinnerFugu Ultra

What I asked each model — the Game prompt

Every model on this page got this exact prompt inside the Agent Operating System: Game — generic 'make a game' open prompt.

Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 24 frontier models have attempted it so far: Claude Fable 5, Fugu Ultra, Fugu Ultra 1.1, Fugu Mini, Fusion, Gemini 3.6 Flash, GLM-5.2, GPT-5.6 Sol, Grok, Inkling, Kimi K3, MiniMax M3, Hermes MoA, Opus 4.8, Claude Opus 5, Qwen 3.8, Qwen 3.7, Claude Sonnet 5, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.7, Kimi K2.7 · Fast, Kimi K2.7 · No-Think, Kimi K2.7 · Quality.

Why this task matters. Game is a textbook test of game-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. A model that ships this in one shot is usually safe to wire into your agent loop for harder tasks of the same shape.

How each model handled Game

Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.

Claude Fable 5 Anthropic
• 7.8/10

What I saw: Clean 3D neon arena renders well with polished glowing walls, floating gold orbs, red chasers, and complete mechanics (score/lives/invuln/restart), but it's a fairly familiar collect-and-dodge concept that lands just short of the field's best.

▶ Play Claude Fable 5's attempt →
Fugu Ultra Sakana AI
🥇 9.0/10 · winner · most reactive

What I saw: Ultra v2 — juicy browser game. Smoke-test PASS with 24.7% pixel diff — one of the most reactive builds on the bench.

▶ Play Fugu Ultra's attempt →
Fugu Ultra 1.1 Sakana AI
• 7.8/10

What I saw: Polished 3D low-poly arena shooter with clean HUD, working radar showing 9 hostiles, HP damage ('ARMOR HIT', HP down to 045) and full combat controls—clearly a functional game, not a walking sim. Held back from top tier by the odd/janky player craft model and no enemies actually visible on-screen in this frame, so combat readability is weaker than the radar suggests.

▶ Play Fugu Ultra 1.1's attempt →
Fugu Mini Sakana AI
🥇 9.0/10 · winner · biggest visual change

What I saw: Juicy browser game. Smoke-test PASS with 55% pixel diff — most reactive build in the sweep.

▶ Play Fugu Mini's attempt →
Fusion OpenRouter
🥇 9.0/10

What I saw: RETRY @ 24K tokens — now complete: 30KB juicy arcade build with 2 rAFs + 9 input handlers + Web Audio + closed tags. Score HUD, lives, polish. The original truncated version has been replaced.

▶ Play Fusion's attempt →
• 8.2/10

What I saw: Strong cyberpunk 3D combat game with polished HUD, working score/kill tracking (250 pts, 1 kill), reduced HP showing active combat, and a well-modeled 3D robot enemy with atmospheric neon lighting; loses points because only one enemy is visible and the scene reads slightly static/sparse compared to the field's best swarm-combat entries.

▶ Play Gemini 3.6 Flash's attempt →
GLM-5.2 Zhipu / Z.ai
• 7.5/10

What I saw: 25KB · plays clean · audio

▶ Play GLM-5.2's attempt →
GPT-5.6 Sol OpenAI
• 8.6/10 · Polished neon arcade

What I saw: Strong: the screenshot renders a genuinely polished neon collect-and-evade game with glowing player, orbiting crystals, pulse-radius ring, HUD stats, energy bar, and clean 'Harvest the Light' title — cohesive and shippable. Weak: it's a familiar collector concept, but the execution (combos, pulse mechanic, lives, best-score persistence, mobile controls) elevates it above the field's best.

▶ Play GPT-5.6 Sol's attempt →
Grok xAI
🥇 9.0/10 · winner · juicy game

What I saw: Open-ended 'make a game' — Grok shipped a juicy 28KB build with score HUD, lives, sound, polish.

▶ Play Grok's attempt →
Inkling Thinking Machines
• 7.4/10

What I saw: Renders cleanly with polished neon 3D visuals, glowing gradient title, grid arena, and functional falling-orb catch mechanic. Strong presentation but gameplay is shallow/generic — no lives, misses, difficulty ramp, or lose condition, so it lands short of the task's best entries.

▶ Play Inkling's attempt →

The winner on Game

Fugu Ultra took gold on this task. winner · most reactive.

What I saw: Ultra v2 — juicy browser game. Smoke-test PASS with 24.7% pixel diff — one of the most reactive builds on the bench.

See Fugu Ultra's full model card: /models/fugu.

Every attempt — live, playable

Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.

How I scored Game — methodology

Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.

My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly