
Real head-to-head · same prompt, one shot
Grok 4.6 vs Hy3
Frontier intelligence at half the frontier price. vs Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval.
Head-to-head verdict: tied 3–3.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Grok 4.6 and Hy3, side by side, on 6 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Grok 4.6 · Benched on 16 skill-infused game builds, every one played through its full gameplay arc before scoring, then the broken ones were handed back to Grok 4.6 to repair itself.
Hy3 · Wired into the Agent OS as the 'Hy3 Coder' tab (chat + live preview + workspace) via OpenRouter. Bench built one-shot on the same prompts as the field; weak builds iterated by Hy3 itself (the model fixes its own builds).
Side-by-side on 21 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Grok 4.6
Hy3
Game
Game
Game
Game
Game
Game
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Game
— not attempted —
Page
— not attempted —
Where Grok 4.6 beat Hy3
The tasks where I gave Grok 4.6 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Doom
Game
Grok 4.6 6.4
·
Hy3 4.5
(+1.9)
What I saw: Re-scored after a real-GPU playtest. It LOOKS like a top-tier Doom - a red demon filling the screen, a shotgun sprite, a working minimap with enemy dots, and ammo that decrements on fire. But at a real 70fps each demon calls hurtPlayer(12) on a 0.18s cooldown, about 67 damage per…
Gtafoot
Game
Grok 4.6 7.4
·
Hy3 7.2
(+0.2)
What I saw: Third-person street level with a real armed character model, long cast shadows, a sunset skybox, pedestrians and parked cars down the block, plus a grid minimap. Played 12s: AMMO went 11 to 09 and SCORE 000010 to 000030 on click-fire, so shooting and scoring are wired. World is b…
Gtadrive
Game
Grok 4.6 7.6
·
Hy3 7.4
(+0.2)
What I saw: Drivable city sandbox: multi-part orange sedan with taillights, lit tower blocks, marked roads, traffic cars, street lamps, a live minimap and a wanted-star row. Played 15s: MPH read 000 then 071 then 016 through accelerate-and-brake, so the driving model integrates properly. Lig…
Where Hy3 beat Grok 4.6
The tasks where I gave Hy3 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Flightsim
Game
Hy3 6.8
·
Grok 4.6 5.0
(+1.8)
What I saw: Clean HUD with working attitude indicator, heading tape, throttle/gear/score panels, and a decently modeled aircraft with wings, nav lights and prop; but the terrain reads as an empty green haze with no visible runway, trees or structures from this altitude, leaving the world fla…
Parachute
Game
Hy3 6.8
·
Grok 4.6 5.4
(+1.4)
What I saw: Renders cleanly with a detailed articulated skydiver (helmet, suit, arms, boots — not a bare capsule), clean HUD with altitude/distance/phase, and a lush jungle canopy of blobs with a visible river below; but the canopy overhead reads as a flat pink slab rather than a parachute, …
Dragonrealm
Game
Hy3 7.2
·
Grok 4.6 6.6
(+0.6)
What I saw: Strong atmospheric snowy world with layered pines, soft shadows, snowfall, and clean HUD (health/stamina/compass/sword chip), but the hero reads as a stubby hooded blob with hidden face and no visible arms/legs, undercutting the flagship Skyrim-ranger fantasy.
Strengths & weaknesses I logged
Grok 4.6
Strengths
- Ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5 Max
- Trained with agentic reinforcement learning for long-running agents, so it holds a spec across a 40KB single-file build
- Fixes its own broken builds: handed the exact runtime error, it repaired 4 of 4 failed games on the first retry
- Very strong arcade and shooter output - the synthwave racer and the Doom raycaster are top-tier one-shots
Trade-offs
- Flight models are its weak spot: both the flight sim and the dogfight shipped unflyable on the first pass
- Two of sixteen game builds died on a hard error (a duplicate identifier and a bad computeBoundingSphere call)
- Worlds are lit and composed but usually untextured, so terrain reads as flat coloured planes
Hy3
Strengths
- Apache-2.0 open weights — self-host free, no lock-in
- Tencent's 270-expert blind eval: 2.67/4 vs GLM-5.1's 2.51, strongest on frontend / data / CI-CD
- Hallucination rate cut 12.5% → 5.4%; stable tool-calls across scaffoldings (<4% SWE-Bench variance)
Trade-offs
- Slow upstream on OpenRouter (30-90s per build) — fine for one-shots, sluggish for tight loops
- One-shot game builds can under-render (flat raycaster walls, unlit 3D) without an iterate pass
Pricing & context — the spec sheet
| Spec | Grok 4.6 | Hy3 |
|---|---|---|
| Vendor | xAI | Tencent Hunyuan |
| Context window | 500,000 tokens | 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. |
| Price | $2 in / $6 out per M tokens | $0.14 / 1M input · $0.58 / 1M output |
| Pricing detail | xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial Analysis Intelligence Index score of 61. | Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but per-token cost is negligible. |
| Release | 2026-08 | 2026-07-06 |
| Bench coverage | 20/20 scored · avg 5.95/10 | 7/7 scored · avg 6.76/10 |
The verdict — which should you pick?
Across 6 scored shared tasks, the averages are essentially tied — Grok 4.6 6.40 vs Hy3 6.65. This isn't the comparison where one wins; it's the comparison where you pick based on context, pricing, and what you're actually trying to ship.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Grok 4.6 and Hy3 both into the Agent Operating System and dispatch each from the kanban by task type — high-volume agent work where the per-token bill decides what you can afford to run → Grok 4.6, cost-sensitive coding + frontend design where open weights matter → Hy3. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Grok 4.6 vs Hy3
Which is better, Grok 4.6 or Hy3?
On Goldie Bench, Grok 4.6 averages 6.40/10 across the shared tasks, with 1 gold, 0 silver, 0 bronze overall. Hy3 averages 6.65/10, with 0 gold, 0 silver, 0 bronze. It's a curated tie on the head-to-head.
How much does Grok 4.6 cost vs Hy3?
Grok 4.6: xAI's August 2026 flagship. $2 per million input tokens and $6 per million output tokens on the standard tier (the Fast tier is double), which is roughly half what the other frontier models charge for the same Artificial Analysis Intelligence Index score of 61. Hy3: Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but per-token cost is negligible.
What's the context window for Grok 4.6 vs Hy3?
Grok 4.6 has a 500,000 tokens context window. Hy3 has a 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. context window.
When should I pick Grok 4.6 over Hy3?
Pick Grok 4.6 for: High-volume agent work where the per-token bill decides what you can afford to run; Arcade, shooter and driving builds in one shot; Self-repair loops - it is unusually good at fixing a build when you hand it the real error. The trade-off is the weaknesses we logged on the bench: Flight models are its weak spot: both the flight sim and the dogfight shipped unflyable on the first pass; Two of sixteen game builds died on a hard error (a duplicate identifier and a bad computeBoundingSphere call); Worlds are lit and composed but usually untextured, so terrain reads as flat coloured planes.
When should I pick Hy3 over Grok 4.6?
Pick Hy3 for: Cost-sensitive coding + frontend design where open weights matter; Self-hosters who want an Apache-2.0 model they fully own; Anyone wiring a cheap capable coder into a live build panel (Agent OS Hy3 Coder tab). The trade-off is the weaknesses we logged on the bench: Slow upstream on OpenRouter (30-90s per build) — fine for one-shots, sluggish for tight loops; One-shot game builds can under-render (flat raycaster walls, unlit 3D) without an iterate pass.
How does Goldie Bench score Grok 4.6 vs Hy3?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Grok 4.6 vs Fusion Hy3 vs Fusion Grok 4.6 vs Claude Opus 5 Hy3 vs Claude Opus 5 Grok 4.6 vs Hermes MoA Hy3 vs Hermes MoA Grok 4.6 vs GPT-5.6 Sol Hy3 vs GPT-5.6 SolFull model pages: Grok 4.6 · Hy3 · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly

























