Grok 4.6
Frontier intelligence at half the frontier price.
Reference benchmarks for Grok 4.6
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Grok 4.6 is honest about what's measured.
What is Grok 4.6?
Grok 4.6 is the xAI frontier model with a 500,000 tokens context window, released 2026-08. Tagline: Frontier intelligence at half the frontier price.. Official source: x.ai/news/grok-4-6.
Pricing detail. xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial Analysis Intelligence Index score of 61.
How I use it inside the Agent OS. Benched on 16 skill-infused game builds, every one played through its full gameplay arc before scoring, then the broken ones were handed back to Grok 4.6 to repair itself.
What I built with Grok 4.6
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Grok 4.6 shipped on the bench: 20 one-shot demos across 500,000 tokens of context. Of those, 20 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5 Max
- Trained with agentic reinforcement learning for long-running agents, so it holds a spec across a 40KB single-file build
- Fixes its own broken builds: handed the exact runtime error, it repaired 4 of 4 failed games on the first retry
- Very strong arcade and shooter output - the synthwave racer and the Doom raycaster are top-tier one-shots
Trade-offs
- Flight models are its weak spot: both the flight sim and the dogfight shipped unflyable on the first pass
- Two of sixteen game builds died on a hard error (a duplicate identifier and a bad computeBoundingSphere call)
- Worlds are lit and composed but usually untextured, so terrain reads as flat coloured planes
Best for
- High-volume agent work where the per-token bill decides what you can afford to run
- Arcade, shooter and driving builds in one shot
- Self-repair loops - it is unusually good at fixing a build when you hand it the real error
Every benchmark — Grok 4.6's full scorecard
All 20 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Grok 4.6 deep dive →.
Every demo by Grok 4.6
20 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Grok 4.6 against every other model
Every head-to-head featuring Grok 4.6. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Grok 4.6 vs Fusion Grok 4.6 vs Claude Opus 5 Grok 4.6 vs Hermes MoA Grok 4.6 vs GPT-5.6 Sol Grok 4.6 vs Claude Fable 5 Grok 4.6 vs Qwen 3.8 Grok 4.6 vs Grok Grok 4.6 vs MiniMax M3 Grok 4.6 vs Fugu Ultra Grok 4.6 vs Kimi K3 Grok 4.6 vs GLM-5.2 Grok 4.6 vs Fugu Mini Grok 4.6 vs Muse Spark 1.2 Grok 4.6 vs Opus 4.8 Grok 4.6 vs Kimi K2.7 Grok 4.6 vs Qwable 5 27B Coder Grok 4.6 vs Gemini 3.6 Flash Grok 4.6 vs Claude Sonnet 5 Grok 4.6 vs Qwen 3.7 Grok 4.6 vs Fugu Ultra 1.1 Grok 4.6 vs Inkling Grok 4.6 vs Agents-A1 Grok 4.6 vs Gemma 4 12B · MLX Grok 4.6 vs Laguna XS 2.1 Grok 4.6 vs Qwythos 9B Grok 4.6 vs LongCat-2.0 Grok 4.6 vs Hy3 Grok 4.6 vs Gemma-4 12B CoderRead more on agentos.guide:
Grok 4.6 — frequently asked
What is Grok 4.6?
Grok 4.6 is xAI's AI model — Frontier intelligence at half the frontier price. It has a 500K tokens context window and was released 2026-08.
How good is Grok 4.6 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 5.95/10 across 20 scored tasks, with 1 gold, 0 silver and 0 bronze medals.
How much does Grok 4.6 cost?
$2 in / $6 out per M tokens. xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial
Where can I see Grok 4.6 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.