Grok
Snappy + real-time — the X-native model.
What is Grok?
Grok is the xAI frontier model with a 256,000 tokens context window, released 2026-04. Tagline: Snappy + real-time — the X-native model.. Official source: grok.com.
Pricing detail. Bundled with X (Twitter) Premium subscription — no per-token bill for end users, no individual API pricing for the chat product.
How I use it inside the Agent OS. Used for real-time content workflows where the model needs current X timeline context. Standalone bench scoring pending.
What I built with Grok
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Grok shipped on the bench: 47 one-shot demos across 256,000 tokens of context. Of those, 43 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Real-time access to X timeline data — unique signal no other model has
- Snappy latency on shorter prompts
- 256K context window keeps pace with the open-weights field
Trade-offs
- 13 demos on the bench but zero have curated 0–10 verdicts yet — currently unranked
- API access is gated behind X Premium, awkward for backend agent loops
Best for
- Workflows that need live X / Twitter context
- Snappy prompts where latency matters
- Researchers comparing X-native models against the rest of the field
Every benchmark — Grok's full scorecard
All 43 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Grok deep dive →.
Every demo by Grok
47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Grok against every other model
Every head-to-head featuring Grok. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Grok vs Fusion Grok vs Claude Opus 5 Grok vs Hermes MoA Grok vs GPT-5.6 Sol Grok vs Claude Fable 5 Grok vs Qwen 3.8 Grok vs MiniMax M3 Grok vs Fugu Ultra Grok vs Kimi K3 Grok vs GLM-5.2 Grok vs Fugu Mini Grok vs Opus 4.8 Grok vs Kimi K2.7 Grok vs Qwable 5 27B Coder Grok vs Gemini 3.6 Flash Grok vs Claude Sonnet 5 Grok vs Qwen 3.7 Grok vs Fugu Ultra 1.1 Grok vs Inkling Grok vs Agents-A1 Grok vs Gemma 4 12B · MLX Grok vs Laguna XS 2.1 Grok vs Qwythos 9B Grok vs LongCat-2.0 Grok vs Hy3 Grok vs Gemma-4 12B CoderRead more on agentos.guide: /grok-build
Grok — frequently asked
What is Grok?
Grok is xAI's AI model — Snappy + real-time — the X-native model. It has a 256K tokens context window and was released 2026-04.
How good is Grok at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 8.09/10 across 43 scored tasks, with 5 gold, 1 silver and 1 bronze medals.
How much does Grok cost?
Subscription via X Premium. Bundled with X (Twitter) Premium subscription — no per-token bill for end users, no individual API pricing for the chat product.
Where can I see Grok demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.