Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

Grok 4.7's full benchmark breakdown.

Grok 4.6's price, a bigger base model, and it builds games that hold together. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-09-22
Scored tasks
11
Reading time
5 min
External sources
4

01 · The headline numbers

GoldieBench average
6.86/10
11 scored one-shot tasks
Board rank
#21
of 24 ranked frontier models
Task medals
0🥇 1🥈 2🥉
outright wins on shared briefs
Context
500,000 tokens
$2 in / $6 out per M tokens

02 · Every benchmark, bar by bar

All 11 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“A gorgeous 3D breakout with rich lighting, gradient sky dome, gridded arena, minimap, and a polished sci-fi HUD; playtest confirms real motion (walkPx 0.23) with hull/lives/score/bricks all changing and no console errors — clearly at or above the field's best.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are Grok 4.7's lowest scored builds, verdicts unedited:

“The 3D world never rendered — the canvas stays black with only the HUD/hotbar visible, and a console error (computeBoundingSphere is not a function) plus zero pixel-change and frozen HUD confirm the voxel scene never appeared. Nicely designed HUD, but the actual game is broken/non-rendering.”

- the judge's verdict, unedited

“The 3D scene never rendered — screen is pure black (lum mean 2.3, zero pixel change across all playtest moments) with a fatal 'hud is not a function' console error freezing the game at DEPTH 000 / SCORE 0000. Only the polished HUD chrome shows; the actual crypt crawler is non-functional.”

- the judge's verdict, unedited

“Strong Doom-flavored raycaster with polished brick walls, weapon sprite, minimap, and detailed demon art in source — the render looks atmospheric and on-brief. But the playtest shows HP draining to 0 (player died) with KILLS stuck at 00 and no monster visible in the frame, so combat against the chasing monsters never actually landed on screen despite the enemies clearly being coded.”

- the judge's verdict, unedited

“Renders a polished HUD, minimap, and a nicely modeled plane over a warm desert/water world, but playtest shows score/kills frozen at 00/08 and 0000 with the game stuck in 'RETURN TO AO' and near-zero late pixel change — the combat loop never engaged. Strong visual presentation undercut by no demonstrable dogfighting.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (11 tasks)

Grok 4.76.9
field avg7.1

Others (0 tasks)

field avg7.5

Pages (0 tasks)

field avg7.6

Sims (0 tasks)

field avg6.9

Visuals (0 tasks)

field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

From the source guides

MeasureResultSource guide
Artificial Analysis Intelligence Index v4.3.246/artificialanalysis.ai
CursorBench 4.046.3%/x.ai
DeepSWE v1.1 (high effort)71.0%/x.ai
GDPval1,695 Elo/x.ai
Terminal-Bench 4.038.0%/x.ai

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does Grok 4.7 perform on GoldieBench?

Grok 4.7 averages 6.86/10 across 11 scored one-shot build tasks, ranking #21 of 24 ranked frontier models, with 0 task golds, 1 silvers and 2 bronzes. Every score is a real judged build you can open and play on this site.

02What is Grok 4.7 best at?

Its strongest scored build is Arcade at 8.6/10. By category it averages Game 6.9 on the bench.

03How much does Grok 4.7 cost?

xAI's September 2026 release, priced exactly like Grok 4.6: $2 per million input tokens and $6 per million output on the standard tier, $4 / $12 on the Fast tier (double the output speed). OpenRouter lists the same model at $1.60 / $4.80.

Source ledger

  1. 01Official vendor site: x.ai/news/grok-4-7x.ai
  2. 02Artificial Analysis Intelligence Index v4.3.2agentos.guide
  3. 03CursorBench 4.0agentos.guide
  4. 04Julian's guide: /grok-4-7-agent-osagentos.guide

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly