Hy3
Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval.
Reference benchmarks for Hy3
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Hy3 is honest about what's measured.
What is Hy3?
Hy3 is the Tencent Hunyuan frontier model with a 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. context window, released 2026-07-06. Tagline: Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval.. Official source: openrouter.ai/tencent/hy3.
Pricing detail. Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but per-token cost is negligible.
How I use it inside the Agent OS. Wired into the Agent OS as the 'Hy3 Coder' tab (chat + live preview + workspace) via OpenRouter. Bench built one-shot on the same prompts as the field; weak builds iterated by Hy3 itself (the model fixes its own builds).
What I built with Hy3
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Hy3 shipped on the bench: 7 one-shot demos across 262,144-token context window. Open weights (Apache-2.0) on HuggingFace / ModelScope / GitHub; benched here via OpenRouter. of context. Of those, 7 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Apache-2.0 open weights — self-host free, no lock-in
- Tencent's 270-expert blind eval: 2.67/4 vs GLM-5.1's 2.51, strongest on frontend / data / CI-CD
- Hallucination rate cut 12.5% → 5.4%; stable tool-calls across scaffoldings (<4% SWE-Bench variance)
Trade-offs
- Slow upstream on OpenRouter (30-90s per build) — fine for one-shots, sluggish for tight loops
- One-shot game builds can under-render (flat raycaster walls, unlit 3D) without an iterate pass
Best for
- Cost-sensitive coding + frontend design where open weights matter
- Self-hosters who want an Apache-2.0 model they fully own
- Anyone wiring a cheap capable coder into a live build panel (Agent OS Hy3 Coder tab)
Every benchmark — Hy3's full scorecard
All 7 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Hy3 deep dive →.
Every demo by Hy3
7 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Hy3 against every other model
Every head-to-head featuring Hy3. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Hy3 vs Fusion Hy3 vs Claude Opus 5 Hy3 vs Hermes MoA Hy3 vs GPT-5.6 Sol Hy3 vs Claude Fable 5 Hy3 vs Qwen 3.8 Hy3 vs Grok Hy3 vs MiniMax M3 Hy3 vs Fugu Ultra Hy3 vs Kimi K3 Hy3 vs GLM-5.2 Hy3 vs Fugu Mini Hy3 vs Muse Spark 1.2 Hy3 vs Opus 4.8 Hy3 vs Kimi K2.7 Hy3 vs Qwable 5 27B Coder Hy3 vs Gemini 3.6 Flash Hy3 vs Claude Sonnet 5 Hy3 vs Qwen 3.7 Hy3 vs Fugu Ultra 1.1 Hy3 vs Inkling Hy3 vs Grok 4.6 Hy3 vs Agents-A1 Hy3 vs Gemma 4 12B · MLX Hy3 vs Laguna XS 2.1 Hy3 vs Qwythos 9B Hy3 vs LongCat-2.0 Hy3 vs Gemma-4 12B CoderRead more on agentos.guide: /hy3-agent-os
Hy3 — frequently asked
What is Hy3?
Hy3 is Tencent Hunyuan's AI model — Tencent's open-weights coder — Apache-2.0, cheap, beats GLM-5.1 on frontend in Tencent's blind eval. It has a 262K tokens context window and was released 2026-07-06.
How good is Hy3 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 6.76/10 across 7 scored tasks, with 0 gold, 0 silver and 0 bronze medals.
How much does Hy3 cost?
$0.14 / 1M input · $0.58 / 1M output. Tencent Hunyuan 3 — open-weights under Apache-2.0, so free to self-host. On OpenRouter it is one of the cheapest capable coders: ~$0.14/M in, $0.58/M out (1 RMB / 4 RMB). Upstream can be slow (30-90s to first token), but
Where can I see Hy3 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.