Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Tencent Hunyuan

Hy3

Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval.

Context262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter.
Pricing$0.14 / 1M input · $0.58 / 1M output
Tasks tested7
Avg score6.76/10 average
Medals🥇0 🥈0 🥉0
Release2026-07-06
Official vendor source
Hy3 is built by Tencent Hunyuan — see the vendor's own product page, pricing, and docs at openrouter.ai/tencent/hy3.
Visit openrouter.ai/tencent/hy3 →

Reference benchmarks for Hy3

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Hy3 is honest about what's measured.

Expert blind eval
2.67/4
Hallucination rate
5.4%

What is Hy3?

Hy3 is the Tencent Hunyuan frontier model with a 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. context window, released 2026-07-06. Tagline: Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval.. Official source: openrouter.ai/tencent/hy3.

Pricing detail. Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but per-token cost is negligible.

How I use it inside the Agent OS. Wired into the Agent OS as the 'Hy3 Coder' tab (chat + live preview + workspace) via OpenRouter. Bench built one-shot on the same prompts as the field; weak builds iterated by Hy3 itself (the model fixes its own builds).

What I built with Hy3

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Hy3 shipped on the bench: 7 one-shot demos across 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. of context. Of those, 7 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Apache-2.0 open weights — self-host free, no lock-in
  • Tencent's 270-expert blind eval: 2.67/4 vs GLM-5.1's 2.51, strongest on frontend / data / CI-CD
  • Hallucination rate cut 12.5% → 5.4%; stable tool-calls across scaffoldings (<4% SWE-Bench variance)

Trade-offs

  • Slow upstream on OpenRouter (30-90s per build) — fine for one-shots, sluggish for tight loops
  • One-shot game builds can under-render (flat raycaster walls, unlit 3D) without an iterate pass

Best for

  • Cost-sensitive coding + frontend design where open weights matter
  • Self-hosters who want an Apache-2.0 model they fully own
  • Anyone wiring a cheap capable coder into a live build panel (Agent OS Hy3 Coder tab)

Every benchmark — Hy3's full scorecard

All 7 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Hy3 deep dive →.

Every demo by Hy3

7 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

Hy3 one-shot build of Dragonrealm — GoldieBench AI benchmark screenshot▶ LIVE
Dragonrealm
Game
Strong atmospheric snowy world with layered pines, soft shadows, snowfall, and clean HUD (health/stamina/compass/sword chip), but the hero reads as a stubby hooded blob with hidden face and no visible arms/legs, undercutting the flagship Skyrim-ranger fantasy.
Hy3 one-shot build of Doom — GoldieBench AI benchmark screenshot▶ LIVE
Doom
Game
HUD, minimap with dot-enemies, and a weapon read clearly, but the main view is a nearly featureless brown wall with no visible walls-vs-open geometry and no on-screen demon sprite, so the raycast world and monsters chasing you don't come across to the player. Solid UI polish can't rescue a flat, empty central frame.
Hy3 one-shot build of Gtafoot — GoldieBench AI benchmark screenshot▶ LIVE
Gtafoot
Game
Strong dusk-city atmosphere with a readable blocky hero (cap, jacket trim, gun), streetlights, crosswalks, and a pedestrian, plus clean HUD (ammo, health bar, wanted stars, crosshair, controls). Weak points: the world is fairly barren mid-frame, no visible buildings-as-cover in this angle, and it reads more as a stylized diorama than the dense GTA street the brief implies, keeping it below the field's best.
Hy3 one-shot build of Gtadrive — GoldieBench AI benchmark screenshot▶ LIVE
Gtadrive
Game
Clean render with a readable yellow hero car (wheels, cabin, taillight), colorful blocky city, working HUD/minimap and speed at 101km/h shows live play. Solid and functional but visually flat-lit and generic — buildings read as bare boxes and lighting is dim, keeping it below the top tier.
Hy3 one-shot build of Aipbpromo — GoldieBench AI benchmark screenshot▶ LIVE
Aipbpromo
Page
Clean multi-scene pipeline with animated count-up stats, gold/cyan cinematic palette, particle field and progress bar render correctly; but the stats appear left-clustered and off-center with only two of four visible mid-animation, feeling sparse rather than composed, keeping it solid-but-generic rather than a task winner.
Hy3 one-shot build of Parachute — GoldieBench AI benchmark screenshot▶ LIVE
Parachute
Game
Renders cleanly with a detailed articulated skydiver (helmet, suit, arms, boots — not a bare capsule), clean HUD with altitude/distance/phase, and a lush jungle canopy of blobs with a visible river below; but the canopy overhead reads as a flat pink slab rather than a parachute, and the world feels more static/decorative than dynamic mid-drop, keeping it below the field's best.
Hy3 one-shot build of Flightsim — GoldieBench AI benchmark screenshot▶ LIVE
Flightsim
Game
Clean HUD with working attitude indicator, heading tape, throttle/gear/score panels, and a decently modeled aircraft with wings, nav lights and prop; but the terrain reads as an empty green haze with no visible runway, trees or structures from this altitude, leaving the world flat and generic compared to the field's best.
every demo, in a grid · click any one to play

Compare Hy3 against every other model

Every head-to-head featuring Hy3. Verdicts shown for scored pairs.

Hy3 vs Fusion
Fusion leads 7–0
Hy3 vs Claude Opus 5
Claude Opus 5 leads 7–0
Hy3 vs Hermes MoA
Hy3 leads 4–3
Hy3 vs GPT-5.6 Sol
GPT-5.6 Sol leads 6–1
Hy3 vs Claude Fable 5
Claude Fable 5 leads 7–0
Hy3 vs Qwen 3.8
Qwen 3.8 leads 6–0
Hy3 vs Grok
Grok leads 5–1
Hy3 vs MiniMax M3
MiniMax M3 leads 7–0
Hy3 vs Fugu Ultra
Tied 1–1
Hy3 vs Kimi K3
Kimi K3 leads 6–1
Hy3 vs GLM-5.2
GLM-5.2 leads 7–0
Hy3 vs Fugu Mini
Fugu Mini leads 2–0
Hy3 vs Muse Spark 1.2
Muse Spark 1.2 leads 7–0
Hy3 vs Opus 4.8
Opus 4.8 leads 5–2
Hy3 vs Kimi K2.7
Kimi K2.7 leads 5–1
Hy3 vs Qwable 5 27B Coder
Qwable 5 27B Coder leads 2–0
Hy3 vs Gemini 3.6 Flash
Gemini 3.6 Flash leads 7–0
Hy3 vs Claude Sonnet 5
Hy3 leads 4–3
Hy3 vs Qwen 3.7
Qwen 3.7 leads 5–2
Hy3 vs Fugu Ultra 1.1
Fugu Ultra 1.1 leads 4–1
Hy3 vs Inkling
Hy3 leads 5–1
Hy3 vs Grok 4.6
Tied 3–3
Hy3 vs Agents-A1
Hy3 leads 2–0
Hy3 vs Gemma 4 12B · MLX
Hy3 leads 2–0
Hy3 vs Laguna XS 2.1
Hy3 leads 2–0
Hy3 vs Qwythos 9B
Hy3 leads 2–0
Hy3 vs LongCat-2.0
LongCat-2.0 leads 1–0
Hy3 vs Gemma-4 12B Coder
2 shared tasks · unscored
Hy3 vs DeepSeek V4 Flash
7 shared tasks · unscored
Hy3 vs DeepSeek V4 Pro
7 shared tasks · unscored
Hy3 vs Kimi K2.7 · Fast
7 shared tasks · unscored
Hy3 vs Kimi K2.7 · No-Think
7 shared tasks · unscored
Hy3 vs Kimi K2.7 · Quality
7 shared tasks · unscored
Hy3 vs Ornith 1.0
2 shared tasks · unscored
Hy3 vs Claude Mythos 5
Reference-only
Hy3 vs Kilo Code
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Hy3 vs Fusion Hy3 vs Claude Opus 5 Hy3 vs Hermes MoA Hy3 vs GPT-5.6 Sol Hy3 vs Claude Fable 5 Hy3 vs Qwen 3.8 Hy3 vs Grok Hy3 vs MiniMax M3 Hy3 vs Fugu Ultra Hy3 vs Kimi K3 Hy3 vs GLM-5.2 Hy3 vs Fugu Mini Hy3 vs Muse Spark 1.2 Hy3 vs Opus 4.8 Hy3 vs Kimi K2.7 Hy3 vs Qwable 5 27B Coder Hy3 vs Gemini 3.6 Flash Hy3 vs Claude Sonnet 5 Hy3 vs Qwen 3.7 Hy3 vs Fugu Ultra 1.1 Hy3 vs Inkling Hy3 vs Grok 4.6 Hy3 vs Agents-A1 Hy3 vs Gemma 4 12B · MLX Hy3 vs Laguna XS 2.1 Hy3 vs Qwythos 9B Hy3 vs LongCat-2.0 Hy3 vs Gemma-4 12B Coder

Read more on agentos.guide: /hy3-agent-os

Hy3 — frequently asked

What is Hy3?

Hy3 is Tencent Hunyuan's AI model — Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval. It has a 262K tokens context window and was released 2026-07-06.

How good is Hy3 at coding and one-shot builds?

On the GoldieBench one-shot build benchmark it averages 6.76/10 across 7 scored tasks, with 0 gold, 0 silver and 0 bronze medals.

How much does Hy3 cost?

$0.14 / 1M input · $0.58 / 1M output. Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but

Where can I see Hy3 demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly