Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
xAI

Grok 4.6

Frontier intelligence at half the frontier price.

Context500,000 tokens
Pricing$2 in / $6 out per M tokens
Tasks tested20
Avg score5.95/10 average
Medals🥇1 🥈0 🥉0
Release2026-08
Official vendor source
Grok 4.6 is built by xAI — see the vendor's own product page, pricing, and docs at x.ai/news/grok-4-6.
Visit x.ai/news/grok-4-6 →

Reference benchmarks for Grok 4.6

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Grok 4.6 is honest about what's measured.

Artificial Analysis Intelligence Index
61
CursorBench v3.2
69.9%
source: /x.ai
DeepSWE v1.1
65.9%
source: /x.ai
GDPVal-AA v2
1753
source: /x.ai

What is Grok 4.6?

Grok 4.6 is the xAI frontier model with a 500,000 tokens context window, released 2026-08. Tagline: Frontier intelligence at half the frontier price.. Official source: x.ai/news/grok-4-6.

Pricing detail. xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial Analysis Intelligence Index score of 61.

How I use it inside the Agent OS. Benched on 16 skill-infused game builds, every one played through its full gameplay arc before scoring, then the broken ones were handed back to Grok 4.6 to repair itself.

What I built with Grok 4.6

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Grok 4.6 shipped on the bench: 20 one-shot demos across 500,000 tokens of context. Of those, 20 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5 Max
  • Trained with agentic reinforcement learning for long-running agents, so it holds a spec across a 40KB single-file build
  • Fixes its own broken builds: handed the exact runtime error, it repaired 4 of 4 failed games on the first retry
  • Very strong arcade and shooter output - the synthwave racer and the Doom raycaster are top-tier one-shots

Trade-offs

  • Flight models are its weak spot: both the flight sim and the dogfight shipped unflyable on the first pass
  • Two of sixteen game builds died on a hard error (a duplicate identifier and a bad computeBoundingSphere call)
  • Worlds are lit and composed but usually untextured, so terrain reads as flat coloured planes

Best for

  • High-volume agent work where the per-token bill decides what you can afford to run
  • Arcade, shooter and driving builds in one shot
  • Self-repair loops - it is unusually good at fixing a build when you hand it the real error

Every benchmark — Grok 4.6's full scorecard

All 20 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Grok 4.6 deep dive →.

Every demo by Grok 4.6

20 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Grok 4.6 one-shot build of Crypt — GoldieBench AI benchmark screenshot▶ LIVE
Crypt
Game
Torch-lit crypt with a warm falloff, a multi-part seeker character, tiled floor slabs, rubble and a room minimap; VITALS, TORCH and RELICS all present. Played 20s: movement and look both respond. But the camera clips hard into a wall on the mid-play frame, the geometry is untextured tan boxes and nothing in the crypt is animated.
Grok 4.6 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
Clean frozen open world: multi-part warrior with helmet, pack and sword, drifting snow, pines, health and stamina bars and a working compass tape. Played 20s: walk and mouse-look both move the camera and the banner state changed from A SHADOW ON THE SNOW to FROST AND FIRE. Marked down because the promised dragon never appears on screen and the terrain is a mostly empty white field.
Grok 4.6 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
Re-scored after a real-GPU playtest. It LOOKS like a top-tier Doom - a red demon filling the screen, a shotgun sprite, a working minimap with enemy dots, and ammo that decrements on fire. But at a real 70fps each demon calls hurtPlayer(12) on a 0.18s cooldown, about 67 damage per second each, so I died 8 times in 6 seconds and never got a fight. My first pass only scored it 8.4 because the headless renderer ran at 2fps and I never reached contact range - the slow harness hid a hard balance failure. Handed Grok 4.6 that exact evidence and it rebalanced its own combat: the live demo now survives a 20-second fight with kills on the board and HP draining gradually.
Grok 4.6 one-shot build of Racing — GoldieBench AI benchmark screenshot▶ LIVE
Racing
Game
Nice multi-part hover-racer model and a real lap minimap with a full circuit drawn on it, but the world underneath is a single flat purple plane with no visible track surface, so you race across an empty field with the route only readable on the mini-map. Played 15s: HULL fell 100 to 0 and the run ended in WRECKED within twenty seconds.
Grok 4.6 one-shot build of Skyrim — GoldieBench AI benchmark screenshot▶ LIVE
Skyrim
Game
Hard failure and the worst of the run. A JavaScript syntax error, Identifier rune has already been declared, kills the whole script, so the canvas stays 0x0, the world never renders and the HUD sits alone on a flat grey field with YOU DIED already showing. Nothing is playable.
Grok 4.6 one-shot build of Twilightvale — GoldieBench AI benchmark screenshot▶ LIVE
Twilightvale
Game
Moody twilight forest with a caped knight carrying a visible sword, purple crystal enemies that close on you, pines, rocks and a live minimap. Played 20s: SHARDS went 0/7 to 1/7, SCORE to 00310 and VIT fell 100 to 006 because the enemies deal damage. Real combat loop, just very pale ground shading.
Grok 4.6 one-shot build of Voxelcraft — GoldieBench AI benchmark screenshot▶ LIVE
Voxelcraft
Game
One-shot score. Hard failure: the screen rendered pure black behind the HUD and the console threw mesh.computeBoundingSphere is not a function inside rebuildType, so no chunk geometry ever reached the scene. Handed Grok 4.6 that one error line and it fixed the call itself. A real-GPU retest then showed the player dying within 8 seconds of spawn, so it got a second round of its own evidence and fixed survivability too - the live demo is a lit voxel world with a multi-part miner, block hotbar and a day cycle that you can actually walk around in.
Grok 4.6 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
Third-person street level with a real armed character model, long cast shadows, a sunset skybox, pedestrians and parked cars down the block, plus a grid minimap. Played 12s: AMMO went 11 to 09 and SCORE 000010 to 000030 on click-fire, so shooting and scoring are wired. World is blocky and the buildings are flat-faced.
Grok 4.6 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
Drivable city sandbox: multi-part orange sedan with taillights, lit tower blocks, marked roads, traffic cars, street lamps, a live minimap and a wanted-star row. Played 15s: MPH read 000 then 071 then 016 through accelerate-and-brake, so the driving model integrates properly. Lighting is flat and the ground plane is an untextured green sheet.
Grok 4.6 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute
Game
One-shot score. Beautiful freefall: multi-part skydiver in a spread, volumetric clouds, speed streaks, terrain far below and a live altimeter counting 2398m down. But SPACE did nothing across a full 30s descent, the SPACE - DEPLOY prompt never cleared and descent stayed pinned at 052 m/s, so the canopy phase was unreachable. Fed that exact evidence back and Grok 4.6 repaired it itself, and the live demo here is its own fix, with the canopy opening and descent dropping to 005 m/s.
Grok 4.6 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim
Game
One-shot score. The world is the best in the run, a low-poly valley with a river, hangar, control tower, marked runway and a sunset sky, plus a proper attitude ball and compass tape. The flight model was dead though: after nine seconds of throttle and five of back-pressure the HUD still read IAS 042KTS, ALT 0000FT, VS +000FPM, unchanged at every sample. Grok 4.6 fixed its own build from that evidence and the live demo now rotates and climbs to 992ft, though it still stalls if you over-pitch.
Grok 4.6 one-shot build of Arcade — GoldieBench AI benchmark screenshot▶ LIVE
Arcade 🥇
Game
Strong: a fully-rendered 3D breakout with atmospheric cityscape, glowing rails, detailed paddle craft, brick grid, active ball with particle trail, and polished HUD (core/velocity/lives/score/wave). Weak: perspective makes the brick field feel small and the 3D angle slightly complicates gameplay readability, but the execution is ambitious and cohesive.
Grok 4.6 one-shot build of Dogfight — GoldieBench AI benchmark screenshot▶ LIVE
Dogfight
Game
One-shot score. Good bones: multi-part fighter, radar with live enemy blips, hull, boost, altitude and knots readouts, and an island world. But the aircraft spawned at ground level and was destroyed on contact, so HULL cycled 095 to 000 to 100 in a respawn loop and ALT never passed 0048 before the SHOT DOWN overlay. Grok 4.6 self-repaired it from that evidence and the live demo spawns airborne at ALT 0210 and holds full hull through a flight.
Grok 4.6 one-shot build of Neonblaster — GoldieBench AI benchmark screenshot▶ LIVE
Neonblaster
Game
Space shooter with a multi-part player ship (glowing engine nacelles, wing pylons), a formation of pink-and-gold enemy fighters, an asteroid field, tracer streaks, an energy-ring gate and a working radar. Played 12s: HULL INTEGRITY dropped 100 to 064 because the enemies actually attack, which is the thing most one-shot shooters forget.
Grok 4.6 one-shot build of Neoncity — GoldieBench AI benchmark screenshot▶ LIVE
Neoncity
Game
Cyberpunk flythrough that does play: SCORE climbed 000015 to 000269 and EXTRACT distance counted down 2457m to 2303m. But it is unforgiving to the point of broken, the mid-play frame is the camera buried inside a building with a big IMPACT banner and HULL down to 056, and the geometry reads as flat cream slabs rather than a neon city.
Grok 4.6 one-shot build of Neonracer — GoldieBench AI benchmark screenshot▶ LIVE
Neonracer
Game
Strong synthwave scene with glowing pink rails, palm/building environment, retro sun, and a clean 3D ship plus polished HUD (armor/speed/boost, score, IMPACT banner). Vapor-trail particle effects aren't clearly visible in this shot, keeping it just shy of the field's best.
Grok 4.6 one-shot build of Nordiccrypt — GoldieBench AI benchmark screenshot▶ LIVE
Nordiccrypt
Game
The HUD renders nicely with themed health/rune meters, minimap frame, and crosshair, but the 3D scene is entirely black — no dungeon, walls, torches, or ruins visible, so the core first-person crawler failed to render.
Grok 4.6 one-shot build of Outrun — GoldieBench AI benchmark screenshot▶ LIVE
Outrun
Game
Best one-shot of the Grok 4.6 run. Full synthwave stack: magenta laser grid, banked road with painted lines and neon kerbs, a multi-part car with a real headlight pool, light-gates, palm rows and a city skyline, plus barrier obstacles you actually dodge. Played 25s: SPEED climbed 126 to 187, DISTANCE 0053m to 0342m and SCORE ticked to 000150 on a real steer input, so the driving loop is live, not decorative.
Grok 4.6 one-shot build of Raycaster — GoldieBench AI benchmark screenshot▶ LIVE
Raycaster
Game
Strong HUD, minimap and a nicely modeled first-person weapon rig, but the actual maze view is barely visible — mostly flat gray walls with an odd tilted geometry above and no sense of navigable corridors, so the core Wolfenstein maze experience doesn't read clearly in the render.
Grok 4.6 one-shot build of Rpg — GoldieBench AI benchmark screenshot▶ LIVE
Rpg
Game
Top-down Ember Vale with broken columns, trees, scattered loot, slime enemies and a dense minimap; HUD carries vitality, stamina, elixirs, relics, gold and renown. Played 15s: VITALITY dropped 100 to 091 from enemy contact. Held back by the camera sitting so high that the hero reads as a few pixels and the palette washing out to pale sage.
every demo, in a grid · click any one to play

Compare Grok 4.6 against every other model

Every head-to-head featuring Grok 4.6. Verdicts shown for scored pairs.

Grok 4.6 vs Fusion
Fusion leads 18–2
Grok 4.6 vs Claude Opus 5
Claude Opus 5 leads 17–3
Grok 4.6 vs Hermes MoA
Hermes MoA leads 15–5
Grok 4.6 vs GPT-5.6 Sol
GPT-5.6 Sol leads 18–2
Grok 4.6 vs Claude Fable 5
Claude Fable 5 leads 17–3
Grok 4.6 vs Qwen 3.8
Qwen 3.8 leads 14–1
Grok 4.6 vs Grok
Grok leads 13–3
Grok 4.6 vs MiniMax M3
MiniMax M3 leads 17–3
Grok 4.6 vs Fugu Ultra
Fugu Ultra leads 12–4
Grok 4.6 vs Kimi K3
Kimi K3 leads 16–3
Grok 4.6 vs GLM-5.2
GLM-5.2 leads 16–4
Grok 4.6 vs Fugu Mini
Fugu Mini leads 9–3
Grok 4.6 vs Muse Spark 1.2
Muse Spark 1.2 leads 14–5
Grok 4.6 vs Opus 4.8
Opus 4.8 leads 15–4
Grok 4.6 vs Kimi K2.7
Kimi K2.7 leads 8–4
Grok 4.6 vs Qwable 5 27B Coder
Qwable 5 27B Coder leads 11–5
Grok 4.6 vs Gemini 3.6 Flash
Gemini 3.6 Flash leads 15–5
Grok 4.6 vs Claude Sonnet 5
Claude Sonnet 5 leads 14–4
Grok 4.6 vs Qwen 3.7
Qwen 3.7 leads 13–6
Grok 4.6 vs Fugu Ultra 1.1
Fugu Ultra 1.1 leads 12–6
Grok 4.6 vs Inkling
Grok 4.6 leads 14–6
Grok 4.6 vs Agents-A1
Grok 4.6 leads 11–5
Grok 4.6 vs Gemma 4 12B · MLX
Grok 4.6 leads 13–3
Grok 4.6 vs Laguna XS 2.1
Grok 4.6 leads 13–3
Grok 4.6 vs Qwythos 9B
Grok 4.6 leads 12–4
Grok 4.6 vs LongCat-2.0
LongCat-2.0 leads 4–0
Grok 4.6 vs Hy3
Tied 3–3
Grok 4.6 vs Gemma-4 12B Coder
Grok 4.6 leads 1–0
Grok 4.6 vs DeepSeek V4 Flash
20 shared tasks · unscored
Grok 4.6 vs DeepSeek V4 Pro
20 shared tasks · unscored
Grok 4.6 vs Kimi K2.7 · Fast
20 shared tasks · unscored
Grok 4.6 vs Kimi K2.7 · No-Think
20 shared tasks · unscored
Grok 4.6 vs Kimi K2.7 · Quality
20 shared tasks · unscored
Grok 4.6 vs Ornith 1.0
16 shared tasks · unscored
Grok 4.6 vs Claude Mythos 5
Reference-only
Grok 4.6 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Grok 4.6 vs Fusion Grok 4.6 vs Claude Opus 5 Grok 4.6 vs Hermes MoA Grok 4.6 vs GPT-5.6 Sol Grok 4.6 vs Claude Fable 5 Grok 4.6 vs Qwen 3.8 Grok 4.6 vs Grok Grok 4.6 vs MiniMax M3 Grok 4.6 vs Fugu Ultra Grok 4.6 vs Kimi K3 Grok 4.6 vs GLM-5.2 Grok 4.6 vs Fugu Mini Grok 4.6 vs Muse Spark 1.2 Grok 4.6 vs Opus 4.8 Grok 4.6 vs Kimi K2.7 Grok 4.6 vs Qwable 5 27B Coder Grok 4.6 vs Gemini 3.6 Flash Grok 4.6 vs Claude Sonnet 5 Grok 4.6 vs Qwen 3.7 Grok 4.6 vs Fugu Ultra 1.1 Grok 4.6 vs Inkling Grok 4.6 vs Agents-A1 Grok 4.6 vs Gemma 4 12B · MLX Grok 4.6 vs Laguna XS 2.1 Grok 4.6 vs Qwythos 9B Grok 4.6 vs LongCat-2.0 Grok 4.6 vs Hy3 Grok 4.6 vs Gemma-4 12B Coder

Read more on agentos.guide:

Grok 4.6 — frequently asked

What is Grok 4.6?

Grok 4.6 is xAI's AI model — Frontier intelligence at half the frontier price. It has a 500K tokens context window and was released 2026-08.

How good is Grok 4.6 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 5.95/10 across 20 scored tasks, with 1 gold, 0 silver and 0 bronze medals.

How much does Grok 4.6 cost?

$2 in / $6 out per M tokens. xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial

Where can I see Grok 4.6 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly