Rpg
RPG — top-down RPG with sprites, combat, inventory.
What I asked each model — the Rpg prompt
Every model on this page got this exact prompt inside the Agent Operating System: RPG — top-down RPG with sprites, combat, inventory.
Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 24 frontier models have attempted it so far: Claude Fable 5, Fugu Ultra, Fugu Ultra 1.1, Fugu Mini, Fusion, Gemini 3.6 Flash, GLM-5.2, GPT-5.6 Sol, Grok, Inkling, Kimi K3, MiniMax M3, Hermes MoA, Opus 4.8, Claude Opus 5, Qwen 3.8, Qwen 3.7, Claude Sonnet 5, Kimi K2.7 · Fast, Kimi K2.7 · No-Think, Kimi K2.7 · Quality, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.7.
Why this task matters. Rpg is a textbook test of game-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. A model that ships this in one shot is usually safe to wire into your agent loop for harder tasks of the same shape.
How each model handled Rpg
Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.
What I saw: Iterated rebuild renders a polished top-down Emberfall RPG: a checkerboard overworld with trees, a dirt path, a lake, scattered loot, a readable pixel hero and a green slime enemy with an HP bar, plus HP/XP/level/ATK and an inventory panel. Move and attack respond (verified) — strong and shippable.
What I saw: Ultra v2 (gap-fill) — top-down RPG with tilemap + NPCs. Smoke-test PASS (1.1% diff).
What I saw: Polished top-down 3D RPG with visible enemies in aggro rings, active combat ('HIT -13' damage numbers, health at 87), pickups, a shrine objective, and a clean HUD showing inventory/kills/hostiles; strong shippable build, only slightly held back by the inventory system being minimal rather than a true managed bag.
What I saw: Top-down RPG. Smoke-test MAYBE-STATIC — character may be present but minimal response to keys.
What I saw: Top-down RPG with tilemap, NPCs, combat, inventory UI. 26KB of game state. Solid hit on the brief.
What I saw: Strong 3D atmospheric arena with visible enemies (glowing-eyed shadow creatures), working combat/kill tracking, minimap, inventory slots and polished HUD — but the task asked for a top-down RPG with sprites, and this delivers a first-person 3D combat scene, so it drifts from the brief and the death overlay dominates the shot.
What I saw: Strong render: cohesive moonlit forest with layered trees/rocks/flowers, wandering wisp enemies, a clean HUD (HP/mana/XP/gold), quest tracker, minimap, and full control legend — clearly on-brief with combat, inventory, and loot systems in source. Slightly generic character sprite and the static intro message hold it just below the top tier, but it's polished and shippable.
What I saw: 35KB top-down RPG with tilemap, walkable terrain, NPCs, combat, HP/MP UI, inventory. Beats Fusion's lighter 26KB attempt on density.
What I saw: Polished UI chrome (title, inventory grid, hint) renders nicely, but the actual game world is empty—no player sprite, enemies, or items visible, and the inventory slots are blank, suggesting the camera/scene or sprite spawn failed. As an RPG with 'sprites, combat, inventory' the core gameplay simply isn't showing.
The winner on Rpg
Grok took gold on this task. winner · top-down RPG.
What I saw: 35KB top-down RPG with tilemap, walkable terrain, NPCs, combat, HP/MP UI, inventory. Beats Fusion's lighter 26KB attempt on density.
See Grok's full model card: /models/grok. Direct head-to-head against the runner-up: Grok vs Claude Opus 5.
Every attempt — live, playable
Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVEHow I scored Rpg — methodology
Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.
My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.
Related
More game benchmarks: all tasks in the Game category · See the best AI model for Rpg · Back to the leaderboard
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.