Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Grok 4.6's full benchmark breakdown.

Frontier intelligence at half the frontier price. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-08-13
Scored tasks
20
Reading time
6 min
External sources
3

01 · The headline numbers

GoldieBench average
5.95/10
20 scored one-shot tasks
Board rank
#23
of 23 ranked frontier models
Task medals
1🥇 0🥈 0🥉
outright wins on shared briefs
Context
500,000 tokens
$2 in / $6 out per M tokens

02 · Every benchmark, bar by bar

All 20 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Strong: a fully-rendered 3D breakout with atmospheric cityscape, glowing rails, detailed paddle craft, brick grid, active ball with particle trail, and polished HUD (core/velocity/lives/score/wave). Weak: perspective makes the brick field feel small and the 3D angle slightly complicates gameplay readability, but the execution is ambitious and cohesive.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Grok 4.6's lowest scored builds, verdicts unedited:

“One-shot score. Hard failure: the screen rendered pure black behind the HUD and the console threw mesh.computeBoundingSphere is not a function inside rebuildType, so no chunk geometry ever reached the scene. Handed Grok 4.6 that one error line and it fixed the call itself. A real-GPU retest then showed the player dying within 8 seconds of spawn, so it got a second round of its own evidence and fixed survivability too - the live demo is a lit voxel world with a multi-part miner, block hotbar and a day cycle that you can actually walk around in.”

- the judge's verdict, unedited

“Hard failure and the worst of the run. A JavaScript syntax error, Identifier rune has already been declared, kills the whole script, so the canvas stays 0x0, the world never renders and the HUD sits alone on a flat grey field with YOU DIED already showing. Nothing is playable.”

- the judge's verdict, unedited

“The HUD renders nicely with themed health/rune meters, minimap frame, and crosshair, but the 3D scene is entirely black — no dungeon, walls, torches, or ruins visible, so the core first-person crawler failed to render.”

- the judge's verdict, unedited

“One-shot score. Good bones: multi-part fighter, radar with live enemy blips, hull, boost, altitude and knots readouts, and an island world. But the aircraft spawned at ground level and was destroyed on contact, so HULL cycled 095 to 000 to 100 in a respawn loop and ALT never passed 0048 before the SHOT DOWN overlay. Grok 4.6 self-repaired it from that evidence and the live demo spawns airborne at ALT 0210 and holds full hull through a flight.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (20 tasks)

Grok 4.66.0
field avg7.1

Others (0 tasks)

field avg7.5

Pages (0 tasks)

field avg7.6

Sims (0 tasks)

field avg6.9

Visuals (0 tasks)

field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

From the source guides

MeasureResultSource guide
Artificial Analysis Intelligence Index61/artificialanalysis.ai
CursorBench v3.269.9%/x.ai
DeepSWE v1.165.9%/x.ai
GDPVal-AA v21753/x.ai

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Grok 4.6 perform on GoldieBench?

Grok 4.6 averages 5.95/10 across 20 scored one-shot build tasks, ranking #23 of 23 ranked frontier models, with 1 task golds, 0 silvers and 0 bronzes. Every score is a real judged build you can open and play on this site.

02What is Grok 4.6 best at?

Its strongest scored build is Arcade at 8.7/10. By category it averages Game 6.0 on the bench.

03How much does Grok 4.6 cost?

xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial Analysis Intelligence Index score of 61.

Source ledger

  1. 01Official vendor site: x.ai/news/grok-4-6x.ai
  2. 02Artificial Analysis Intelligence Indexagentos.guide
  3. 03CursorBench v3.2agentos.guide

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly