Grok 4.7
Grok 4.6's price, a bigger base model, and it builds games that hold together.
Reference benchmarks for Grok 4.7
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Grok 4.7 is honest about what's measured.
What is Grok 4.7?
Grok 4.7 is the xAI frontier model with a 500,000 tokens context window, released 2026-09. Tagline: Grok 4.6's price, a bigger base model, and it builds games that hold together.. Official source: x.ai/news/grok-4-7.
Pricing detail. xAI's September 2026 release, priced exactly like Grok 4.6: $2 per million input tokens and $6 per million output on the standard tier, $4 / $12 on the Fast tier (double the output speed). OpenRouter lists the same model at $1.60 / $4.80.
How I use it inside the Agent OS. Wired into the Agent OS Grok Build tab (4.7 / 4.6 / 4.5 picker, OpenRouter fallback when the CLI is signed out) and benched on twenty skill-infused game builds, each played through its full gameplay arc on a Metal GPU before the vision judge scored the played frame.
What I built with Grok 4.7
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Grok 4.7 shipped on the bench: 11 one-shot demos across 500,000 tokens of context. Of those, 11 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Averaged 6.86/10 on the same twenty skill-infused game briefs where Grok 4.6 averaged 5.63, winning 8 of the 11 head-to-heads
- Best one-shots on this run: arcade (8.6), dragonrealm (8.6), flightsim (8.5), every one played on a real GPU before scoring
- Longer reinforcement learning shows: it holds a 40-50KB single-file spec and wires the controls it advertises
- xAI's own coding numbers moved: CursorBench 4.0 46.3% (from 40.4%) and DeepSWE v1.1 71.0% (from 65.2%)
Trade-offs
- 2 of 20 builds died on load or never moved when played — crypt: hud is not a function; voxelcraft: mesh.computeBoundingSphere is not a function file:///Users/j
- Independent Artificial Analysis Intelligence Index v4.3.2 puts it at 46, seven points behind Claude Fable 5.1 and GPT-6 at 53
- Terminal-Bench 4.0 is the weak spot xAI reports itself: 38.0% at launch
Best for
- One-shot arcade, driving and shooter builds where the whole game has to arrive in a single file
- High-volume agent work priced at half what the other frontier models charge
- Grok Build inside the Agent OS: pick 4.7 in the tab and everything it writes lands in the workspace
Every benchmark — Grok 4.7's full scorecard
All 11 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Grok 4.7 deep dive →.
Every demo by Grok 4.7
11 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Grok 4.7 against every other model
Every head-to-head featuring Grok 4.7. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Grok 4.7 vs Fusion Grok 4.7 vs Claude Opus 5 Grok 4.7 vs Hermes MoA Grok 4.7 vs GPT-5.6 Sol Grok 4.7 vs Claude Fable 5 Grok 4.7 vs Qwen 3.8 Grok 4.7 vs Grok Grok 4.7 vs MiniMax M3 Grok 4.7 vs Fugu Ultra Grok 4.7 vs Kimi K3 Grok 4.7 vs GLM-5.2 Grok 4.7 vs Fugu Mini Grok 4.7 vs Muse Spark 1.2 Grok 4.7 vs Opus 4.8 Grok 4.7 vs Kimi K2.7 Grok 4.7 vs Qwable 5 27B Coder Grok 4.7 vs Gemini 3.6 Flash Grok 4.7 vs Claude Sonnet 5 Grok 4.7 vs Qwen 3.7 Grok 4.7 vs Fugu Ultra 1.1 Grok 4.7 vs Inkling Grok 4.7 vs Grok 4.6 Grok 4.7 vs Agents-A1 Grok 4.7 vs Gemma 4 12B · MLX Grok 4.7 vs Laguna XS 2.1 Grok 4.7 vs Qwythos 9B Grok 4.7 vs LongCat-2.0 Grok 4.7 vs Hy3 Grok 4.7 vs Gemma-4 12B CoderRead more on agentos.guide: /grok-4-7-agent-os
Grok 4.7 — frequently asked
What is Grok 4.7?
Grok 4.7 is xAI's AI model — Grok 4.6's price, a bigger base model, and it builds games that hold together. It has a 500K tokens context window and was released 2026-09.
How good is Grok 4.7 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 6.86/10 across 11 scored tasks, with 0 gold, 1 silver and 2 bronze medals.
How much does Grok 4.7 cost?
$2 in / $6 out per M tokens. xAI's September 2026 release, priced exactly like Grok 4.6: $2 per million input tokens and $6 per million output on the standard tier, $4 / $12 on the Fast tier (double the output speed). OpenRouter lists the same model
Where can I see Grok 4.7 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.