Game
Game — generic 'make a game' open prompt.
What I asked each model — the Game prompt
Every model on this page got this exact prompt inside the Agent Operating System: Game — generic 'make a game' open prompt.
Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 24 frontier models have attempted it so far: Claude Fable 5, Fugu Ultra, Fugu Ultra 1.1, Fugu Mini, Fusion, Gemini 3.6 Flash, GLM-5.2, GPT-5.6 Sol, Grok, Inkling, Kimi K3, MiniMax M3, Hermes MoA, Opus 4.8, Claude Opus 5, Qwen 3.8, Qwen 3.7, Claude Sonnet 5, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.7, Kimi K2.7 · Fast, Kimi K2.7 · No-Think, Kimi K2.7 · Quality.
Why this task matters. Game is a textbook test of game-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. A model that ships this in one shot is usually safe to wire into your agent loop for harder tasks of the same shape.
How each model handled Game
Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.
What I saw: Clean 3D neon arena renders well with polished glowing walls, floating gold orbs, red chasers, and complete mechanics (score/lives/invuln/restart), but it's a fairly familiar collect-and-dodge concept that lands just short of the field's best.
What I saw: Ultra v2 — juicy browser game. Smoke-test PASS with 24.7% pixel diff — one of the most reactive builds on the bench.
What I saw: Polished 3D low-poly arena shooter with clean HUD, working radar showing 9 hostiles, HP damage ('ARMOR HIT', HP down to 045) and full combat controls—clearly a functional game, not a walking sim. Held back from top tier by the odd/janky player craft model and no enemies actually visible on-screen in this frame, so combat readability is weaker than the radar suggests.
What I saw: Juicy browser game. Smoke-test PASS with 55% pixel diff — most reactive build in the sweep.
What I saw: RETRY @ 24K tokens — now complete: 30KB juicy arcade build with 2 rAFs + 9 input handlers + Web Audio + closed tags. Score HUD, lives, polish. The original truncated version has been replaced.
What I saw: Strong cyberpunk 3D combat game with polished HUD, working score/kill tracking (250 pts, 1 kill), reduced HP showing active combat, and a well-modeled 3D robot enemy with atmospheric neon lighting; loses points because only one enemy is visible and the scene reads slightly static/sparse compared to the field's best swarm-combat entries.
What I saw: Strong: the screenshot renders a genuinely polished neon collect-and-evade game with glowing player, orbiting crystals, pulse-radius ring, HUD stats, energy bar, and clean 'Harvest the Light' title — cohesive and shippable. Weak: it's a familiar collector concept, but the execution (combos, pulse mechanic, lives, best-score persistence, mobile controls) elevates it above the field's best.
What I saw: Open-ended 'make a game' — Grok shipped a juicy 28KB build with score HUD, lives, sound, polish.
What I saw: Renders cleanly with polished neon 3D visuals, glowing gradient title, grid arena, and functional falling-orb catch mechanic. Strong presentation but gameplay is shallow/generic — no lives, misses, difficulty ramp, or lose condition, so it lands short of the task's best entries.
The winner on Game
Fugu Ultra took gold on this task. winner · most reactive.
What I saw: Ultra v2 — juicy browser game. Smoke-test PASS with 24.7% pixel diff — one of the most reactive builds on the bench.
See Fugu Ultra's full model card: /models/fugu.
Every attempt — live, playable
Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVEHow I scored Game — methodology
Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.
My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.
Related
More game benchmarks: all tasks in the Game category · See the best AI model for Game · Back to the leaderboard
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.