Pathtracer
Path Tracer — physically-correct ray-traced renderer.
What I asked each model — the Pathtracer prompt
Every model on this page got this exact prompt inside the Agent Operating System: Path Tracer — physically-correct ray-traced renderer.
Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 23 frontier models have attempted it so far: Claude Fable 5, Fugu Ultra, Fugu Mini, Fusion, Gemini 3.6 Flash, GLM-5.2, GPT-5.6 Sol, Grok, Inkling, Kimi K3, MiniMax M3, Hermes MoA, Opus 4.8, Claude Opus 5, Qwen 3.8, Qwen 3.7, Claude Sonnet 5, Kimi K2.7 · Fast, Kimi K2.7 · No-Think, Kimi K2.7 · Quality, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.7.
Why this task matters. Pathtracer is a textbook test of sim-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. Shipping this cleanly is the floor for what I expect from a frontier model — every model on the leaderboard should at least attempt it.
How each model handled Pathtracer
Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.
What I saw: Textbook Cornell box with convincing colored-wall bleeding, a golden metal sphere with sharp reflections showing the room, and a glass sphere with proper refraction/caustics and a purple sphere visible through it — physically-correct GI, soft shadows and checker floor all render cleanly. Minor Monte-Carlo noise at 80 samples is expected and it converges; a top-tier build on this task.
What I saw: Ultra v2 — WebGL path tracer with sample accumulation. Smoke-test PASS (4.1% pixel diff).
What I saw: Mini gap-fill (round 2) — WebGL path tracer. Smoke-test PASS (0.8% diff).
What I saw: WebGL fragment-shader path tracer — Cornell-box-style scene with accumulating samples, soft shadows, indirect bounce. Real renderer, not faked.
What I saw: Polished UI with full controls (bounces, roughness, DOF, focal, presets) and a plausible path-tracing shader in source, but the render is essentially broken — a flat blue-gray plane with a triangular noise wedge and no visible Cornell box, spheres, or lighting, indicating a camera/geometry or accumulation failure.
What I saw: Real GLSL path tracer with diffuse/metal/glass materials, area light importance sampling and progressive accumulation is technically legit, and the diffuse pink sphere plus glass sphere read convincingly. But the render is visually broken: the floor is largely black with jagged noisy artifacts, the checker plane and refractive spheres show distracting aliasing/fireflies, and the framing feels off — polished chrome but flawed output keeps it below the shippable bar.
What I saw: WebGL fragment-shader path tracer with Cornell-box scene, soft shadows, sample accumulation. 18KB.
What I saw: Genuine Monte Carlo path tracer with cosine-weighted sampling and accumulation, and the neon title/UI is polished — but the render is dominated by a blown-out overexposed emissive sphere that washes over everything, the other spheres are barely visible, and the tone-mapping (sqrt without proper 255 scaling clamp) looks broken, making the scene read as a poor image rather than a convincing render.
What I saw: Genuine GLSL Monte Carlo path tracer rendering a proper Cornell box with correct colored-wall bleeding, metal/glass/diffuse spheres, soft shadows and a refracting glass sphere with visible caustics — clearly physically-based and progressively converging (samples counter live). Only weakness is the noisy, under-converged framebuffer and the header text overlapping a reflected sphere, but the physics and polish edge it past the field's best.
The winner on Pathtracer
Claude Fable 5 took gold on this task. Cornell box GI.
What I saw: Textbook Cornell box with convincing colored-wall bleeding, a golden metal sphere with sharp reflections showing the room, and a glass sphere with proper refraction/caustics and a purple sphere visible through it — physically-correct GI, soft shadows and checker floor all render cleanly. Minor Monte-Carlo noise at 80 samples is expected and it converges; a top-tier build on this task.
See Claude Fable 5's full model card: /models/fable-5.
Every attempt — live, playable
Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVEHow I scored Pathtracer — methodology
Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.
My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.
Related
More sim benchmarks: all tasks in the Sim category · See the best AI model for Pathtracer · Back to the leaderboard
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.