
Real head-to-head · same prompt, one shot
Kimi K3 vs Opus 4.8
Moonshot's 2.8T flagship — 1M context, tuned for long-horizon agent work. vs The reasoning king — deepest thinking, premium price.
Head-to-head verdict: Kimi K3 wins 33–13 with 1 tie.
What I tested — same prompt, two models
I run the same fixed prompt set through every new model the day it drops — same string, one shot, single HTML file out — and I score the result 0–10 on whether it ran, how close it hit the brief, and how good it looked. Below is what came out when I gave the exact same prompts to Kimi K3 and Opus 4.8, side by side, on 47 shared tasks inside the Agent Operating System.
Both models were given identical prompts inside the Agent Operating System — no help, no iteration, no "best of N" tricks. I run each prompt once, save the HTML file the model produces, and score it 0–10 on whether it ran, how close it hit the brief, and how good it looked. The scoring is mine. The verdicts below are pulled from my source comparison guides at agentos.guide where I publish every score and the reasoning behind it.
Kimi K3 · Wired into the Agent OS as the `kimi-k3` Hermes profile and a K3 speed-toggle in the Kimi Code tab — used for long unattended agent runs where a slow-but-right model beats a fast-but-forgetful one.
Opus 4.8 · The default when the build has to ship on the first prompt — Opus is the safety net inside Agent OS for hard one-shots.
Side-by-side on 50 shared tasks
Click any cell to play that model's actual one-shot attempt. Medals are derived from my 0–10 scores per task (highest = 🥇, second = 🥈, third = 🥉).
Task ↓
Kimi K3
Opus 4.8
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Game
Page
Where Kimi K3 beat Opus 4.8
The tasks where I gave Kimi K3 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Webos
Page
Kimi K3 8.6
·
Opus 4.8 5.5
(+3.1)
· polished nebula desktop
What I saw: Strong, highly polished render: crisp macOS-style traffic-light windows, blurred glass dock with running dots, desktop icons, animated 3D wireframe backdrop and starfield, plus a clean welcome/about card — clearly on-brief with Notes/Paint/Terminal. Source confirms real window ma…
Aurora
Visual
Kimi K3 8.7
·
Opus 4.8 6.0
(+2.7)
· volumetric 3D aurora
What I saw: Gorgeous flowing volumetric aurora ribbons with convincing fbm noise, layered mountains, spruce silhouettes, moon, stars and a shooting star make a genuinely atmospheric scene; the elegant typography, palette switcher and vignette give it a shippable polish that edges past the fi…
Nordiccrypt
Game
Kimi K3 8.6
·
Opus 4.8 6.0
(+2.6)
· atmospheric torch-lit crypt
What I saw: Strong: genuine first-person 3D dungeon with convincing torch glow, warm falloff lighting on stone brickwork, scattered rubble props, and a polished serif Nordic UI with rune-collection goal that fully nails the brief. Minor weakness: the corridor geometry is a bit boxy and props…
Pool
Game
Kimi K3 8.1
·
Opus 4.8 5.5
(+2.6)
What I saw: Renders a polished 3D pool table with proper rack, numbered balls, cue stick, aim line, pockets, HUD and physics/audio scaffolding — clearly shippable and on-brief. Weak points: the floating lamp shade looks unlit/misplaced and the rack triangle is slightly loose, keeping it just…
Fluid
Sim
Kimi K3 9.0
·
Opus 4.8 7.0
(+2.0)
What I saw: Gorgeous, textbook-quality WebGL fluid sim with rich swirling dye, added particle sparkle, and polished UI (gradient title, hint pill, control buttons) — the vorticity/pressure-solve pipeline and half-float fallback handling are all correct and shippable; only minor knock is the …
Where Opus 4.8 beat Kimi K3
The tasks where I gave Opus 4.8 a higher 0–10 score on the same prompt — with the actual commentary from my source guides.
Orbit
Sim
Opus 4.8 9.0
·
Kimi K3 3.0
(+6.0)
· winner · accuracy
What I saw: Opus nailed the brief — labelled planet orbits, a real NEO / close-pass panel, a sim clock. GLM went for drama: a glowing nebula swirl that's gorgeous but reads more galaxy than orbit map. Kimi's is accurate but dim and sparse.
Galaxy
Sim
Opus 4.8 8.5
·
Kimi K3 3.0
(+5.5)
· winner · interactive 3D
What I saw: Opus built a proper interactive 3D galaxy — drag to orbit a 7,000-star cloud around a glowing core. Kimi's is the prettiest single frame: a clean tilted spiral disk with rainbow arms. GLM's runs on a canvas with a slick NGC-style HUD and zoom, just less dramatic at a glance. Thre…
Dogfight
Game
Opus 4.8 7.5
·
Kimi K3 3.5
(+4.0)
What I saw: 14KB · plays clean · webgl, input
Terrain
Visual
Opus 4.8 7.0
·
Kimi K3 4.5
(+2.5)
What I saw: 2KB · plays clean · rAF
Solar
Sim
Opus 4.8 8.5
·
Kimi K3 6.5
(+2.0)
· winner · 3D depth
What I saw: Three genuinely good space sims. Opus tilts the orbits into real 3D with a bloom-heavy sun and Saturn's rings. GLM's is the most product-like — labelled planets, orbit and label toggles, a clean HUD. Kimi's is a tidy tilted-orbit system with rings and a deep starfield. Opus and G…
Strengths & weaknesses I logged
Kimi K3
Strengths
- Launch-day benchmarks put it around the Fable/Sol tier, with Terminal Bench (agentic terminal-driving) the standout
- 1M-token context verified on this bench's needle test: exact recall from 162k tokens of noise in 18s
- One-shot builds run long but land complete — its first bench game (13.4 min of thinking, 30,880 tokens) playtested with zero JS errors
- Included in the Kimi coding plan — frontier tier without a new bill
Trade-offs
- Slow on hard tasks — early testers report up to ~35 minutes at max reasoning; this bench saw 13+ minute single builds
- Launch-day rate limits on OpenRouter (429s) — the coding-plan endpoint was the reliable route
- Self-reports as K2.7 if you ask it — verify the served model via the API response, not the model's word
Opus 4.8
Strengths
- Most consistent across the Goldie Bench bench — no weak build, 8.46/10 average
- Deepest one-shot reasoning, especially on game-feel and physics
- Extended thinking mode handles up to 1M tokens of context
Trade-offs
- 5–10× the per-token cost of every other model on the bench
- Less flair on cinematic visuals than GLM-5.2 — playing it safer wins on accuracy, costs you on showpiece moments
Pricing & context — the spec sheet
| Spec | Kimi K3 | Opus 4.8 |
|---|---|---|
| Vendor | Moonshot AI | Anthropic |
| Context window | 1,048,576 tokens — a full codebase in working memory | 200,000 tokens (1M with extended thinking) |
| Price | $3 / M in | $15 / $75 per M tokens |
| Pricing detail | Launched July 16, 2026. 2.8T-param MoE (Moonshot's quickstart corrected the circulating 2.5T estimate). $3/M input on OpenRouter at launch; included at no extra cost in the Kimi coding plan (`k3` on the coding endpoint). | Premium pricing via the Anthropic API: $15 per million input tokens, $75 per million output tokens. Extended thinking is included but adds latency. |
| Release | 2026-07-16 | 2026-05 |
| Bench coverage | 50/50 scored · avg 7.89/10 | 47/47 scored · avg 7.51/10 |
The verdict — which should you pick?
Across 47 scored shared tasks, Kimi K3 averaged 7.86/10, beating Opus 4.8's 7.51/10 by 0.35 points. Pick Kimi K3 when the build has to ship on the first prompt and you can afford the trade-offs in the comparison below.
If you only run one of these inside your stack, the head-to-head average above is the call. If you can run both, my honest play is to wire Kimi K3 and Opus 4.8 both into the Agent Operating System and dispatch each from the kanban by task type — long-horizon agent runs → Kimi K3, mission-critical one-shot builds where 'has to work the first time' matters → Opus 4.8. That's the same setup I run for the 4,000+ founders inside the AI Profit Boardroom.
FAQ — Kimi K3 vs Opus 4.8
Which is better, Kimi K3 or Opus 4.8?
On Goldie Bench, Kimi K3 averages 7.86/10 across the shared tasks, with 8 gold, 4 silver, 7 bronze overall. Opus 4.8 averages 7.51/10, with 3 gold, 1 silver, 0 bronze. Kimi K3 wins the head-to-head 33–13.
How much does Kimi K3 cost vs Opus 4.8?
Kimi K3: Launched July 16, 2026. 2.8T-param MoE (Moonshot's quickstart corrected the circulating 2.5T estimate). $3/M input on OpenRouter at launch; included at no extra cost in the Kimi coding plan (`k3` on the coding endpoint). Opus 4.8: Premium pricing via the Anthropic API: $15 per million input tokens, $75 per million output tokens. Extended thinking is included but adds latency.
What's the context window for Kimi K3 vs Opus 4.8?
Kimi K3 has a 1,048,576 tokens — a full codebase in working memory context window. Opus 4.8 has a 200,000 tokens (1M with extended thinking) context window.
When should I pick Kimi K3 over Opus 4.8?
Pick Kimi K3 for: long-horizon agent runs; whole-repo context work; terminal-driving agents. The trade-off is the weaknesses we logged on the bench: Slow on hard tasks — early testers report up to ~35 minutes at max reasoning; this bench saw 13+ minute single builds; Launch-day rate limits on OpenRouter (429s) — the coding-plan endpoint was the reliable route; Self-reports as K2.7 if you ask it — verify the served model via the API response, not the model's word.
When should I pick Opus 4.8 over Kimi K3?
Pick Opus 4.8 for: Mission-critical one-shot builds where 'has to work the first time' matters; Hard reasoning tasks (planning, multi-step) where you'll pay for the depth; Anything where vendor reliability beats the per-token bill. The trade-off is the weaknesses we logged on the bench: 5–10× the per-token cost of every other model on the bench; Less flair on cinematic visuals than GLM-5.2 — playing it safer wins on accuracy, costs you on showpiece moments.
How does Goldie Bench score Kimi K3 vs Opus 4.8?
Every demo on this page was built by Julian Goldie inside the Agent Operating System — same fixed prompt for both models, one shot, single HTML file out. Each result gets a 0–10 score on whether it ran, how close it hit the brief, and how good it looked. The highest score on each task gets gold; second gets silver; third gets bronze. See methodology for full provenance.
Related comparisons
Other head-to-heads using the same scoring system:
Kimi K3 vs Fusion Opus 4.8 vs Fusion Kimi K3 vs Claude Opus 5 Opus 4.8 vs Claude Opus 5 Kimi K3 vs Hermes MoA Opus 4.8 vs Hermes MoA Kimi K3 vs GPT-5.6 Sol Opus 4.8 vs GPT-5.6 SolFull model pages: Kimi K3 · Opus 4.8 · back to the leaderboard
The same stack Julian uses
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.
4,000+founders
258documented wins
38countries
$59/momonthly














































