Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Visual

Aurora

Aurora — northern lights animation.

CategoryVisual
Models tested24
Scored18/24
Avg score7.38/10
WinnerKimi K3

What I asked each model — the Aurora prompt

Every model on this page got this exact prompt inside the Agent Operating System: Aurora — northern lights animation.

Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 24 frontier models have attempted it so far: Claude Fable 5, Fugu Ultra, Fugu Ultra 1.1, Fugu Mini, Fusion, Gemini 3.6 Flash, GLM-5.2, GPT-5.6 Sol, Grok, Inkling, Kimi K3, MiniMax M3, Hermes MoA, Opus 4.8, Claude Opus 5, Qwen 3.8, Qwen 3.7, Claude Sonnet 5, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.7, Kimi K2.7 · Fast, Kimi K2.7 · No-Think, Kimi K2.7 · Quality.

Why this task matters. Aurora is a textbook test of visual-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. Shipping this cleanly is the floor for what I expect from a frontier model — every model on the leaderboard should at least attempt it.

How each model handled Aurora

Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.

Claude Fable 5 Anthropic
🥈 8.6/10 · shader aurora curtains

What I saw: Beautiful WebGL fragment shader with layered green-blue aurora ribbons, twinkling starfield, snow-rimmed mountain silhouette, and elegant title typography — genuinely convincing northern lights. Interactive mouse sway and click surge plus the reflected glow push it to task-topping polish.

▶ Play Claude Fable 5's attempt →
Fugu Ultra Sakana AI
• 8.0/10

What I saw: Ultra v2 (gap-fill) — aurora over mountain ridge. Smoke-test PASS (0.8% diff — visual-only prompt).

▶ Play Fugu Ultra's attempt →
Fugu Ultra 1.1 Sakana AI
• 5.5/10

What I saw: Clean 3D scene with moon, starfield, layered mountains and pyramid 'trees', but the actual aurora is barely visible—only faint green ground-glow blobs rather than the rippling northern-lights sky the brief demands. Polished framing and typography, but it fails the core aurora effect.

▶ Play Fugu Ultra 1.1's attempt →
Fugu Mini Sakana AI
• 6.5/10

What I saw: Aurora ribbons over mountain ridge. Smoke-test MAYBE-STATIC (<0.5% pixel diff after input) — animates but doesn't respond to keys (which is expected for this visual-only prompt).

▶ Play Fugu Mini's attempt →
Fusion OpenRouter
• 7.5/10

What I saw: Aurora ribbons over a mountain ridge silhouette, WebGL noise shader. Slow morphing curtains in cyan + magenta. Smaller scope (10KB) but the prompt is met.

▶ Play Fusion's attempt →
• 8.4/10 · 3D aurora curtains

What I saw: Strong WebGL aurora with convincing vertical curtain striations, glowing lower edge, and a reflective terrain foreground; polished glass UI with palette/speed/intensity controls. Slightly loses the top of the field because the color gradient bands could read more organically and the terrain grid looks a touch artificial, but overall shippable and near the best.

▶ Play Gemini 3.6 Flash's attempt →
GLM-5.2 Zhipu / Z.ai
• 7.0/10

What I saw: 8KB · plays clean · plain

▶ Play GLM-5.2's attempt →
GPT-5.6 Sol OpenAI
🥈 8.6/10 · Cinematic aurora scene

What I saw: Gorgeous layered green-to-violet curtains with soft blur, twinkling stars, silhouetted mountains and elegant typography make this genuinely cinematic and on-brief. Interactive wind/tap hints and Kp status polish it; only minor risk is the aurora ribbons overlapping the H1 slightly, but overall it matches the field's best.

▶ Play GPT-5.6 Sol's attempt →
Grok xAI
• 7.0/10

What I saw: Aurora ribbons over mountain ridge with stars. Simpler build (9KB) — lighter on detail than Fusion.

▶ Play Grok's attempt →
Inkling Thinking Machines
• 7.8/10

What I saw: Renders a vivid, colorful WebGL aurora curtain with clean green/cyan/purple bands, ground plane, and tasteful title overlay — clearly on-brief and polished; but the curtain reads as a contained rectangular slab rather than a sky-spanning flowing veil, and the mountain silhouette is invisible, keeping it just short of the field's best.

▶ Play Inkling's attempt →

The winner on Aurora

Kimi K3 took gold on this task. volumetric 3D aurora.

What I saw: Gorgeous flowing volumetric aurora ribbons with convincing fbm noise, layered mountains, spruce silhouettes, moon, stars and a shooting star make a genuinely atmospheric scene; the elegant typography, palette switcher and vignette give it a shippable polish that edges past the field's best.

See Kimi K3's full model card: /models/kimik3. Direct head-to-head against the runner-up: Kimi K3 vs Claude Fable 5.

Every attempt — live, playable

Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.

How I scored Aurora — methodology

Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.

My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly