Kimi K2.7
The heavy lifter — frontier coder at flat-rate.
Reference benchmarks for Kimi K2.7
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Kimi K2.7 is honest about what's measured.
What is Kimi K2.7?
Kimi K2.7 is the Moonshot AI frontier model with a 256,000 tokens context window, released 2026-06. Tagline: The heavy lifter — frontier coder at flat-rate.. Official source: kimi.com.
Pricing detail. Available on Moonshot's flat-rate subscription plan — no per-token billing for individual builders. The plan covers all three speed modes (Fast, No-Think, Quality). Vendor: Moonshot AI (moonshot.ai), based in Beijing.
How I use it inside the Agent OS. Wired into the Agent OS as the heavy-lifter for game/sim prototypes and Kanban-dispatched code work. Mode toggled per task: Quality for one-shot games, Fast for short bursts.
What I built with Kimi K2.7
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Kimi K2.7 shipped on the bench: 47 one-shot demos across 256,000 tokens of context. Of those, 25 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Best-of-three on interactive games — raycaster, DOOM, monster AI
- Three speed modes (Fast / No-Think / Quality) you can swap per task
- Flat-rate plan eliminates the per-token meter, so iteration is free
Trade-offs
- Plays plainest on abstract visual prompts — synthwave grids, fluid sims, aurora — where GLM and Opus add more flair
- Bronze average on the Goldie Bench bench despite the gold-medal games — its visual builds are accurate but understated
Best for
- Interactive game prototypes you want shippable on the first prompt
- High-iteration agent loops where per-token cost would dominate
- Long-context refactors using the 256K window inside Agent OS
Every benchmark — Kimi K2.7's full scorecard
All 25 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Kimi K2.7 deep dive →.
Every demo by Kimi K2.7
47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Kimi K2.7 against every other model
Every head-to-head featuring Kimi K2.7. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Kimi K2.7 vs Fusion Kimi K2.7 vs Claude Opus 5 Kimi K2.7 vs Hermes MoA Kimi K2.7 vs GPT-5.6 Sol Kimi K2.7 vs Claude Fable 5 Kimi K2.7 vs Qwen 3.8 Kimi K2.7 vs Grok Kimi K2.7 vs MiniMax M3 Kimi K2.7 vs Fugu Ultra Kimi K2.7 vs Kimi K3 Kimi K2.7 vs GLM-5.2 Kimi K2.7 vs Fugu Mini Kimi K2.7 vs Opus 4.8 Kimi K2.7 vs Qwable 5 27B Coder Kimi K2.7 vs Gemini 3.6 Flash Kimi K2.7 vs Claude Sonnet 5 Kimi K2.7 vs Qwen 3.7 Kimi K2.7 vs Fugu Ultra 1.1 Kimi K2.7 vs Inkling Kimi K2.7 vs Agents-A1 Kimi K2.7 vs Gemma 4 12B · MLX Kimi K2.7 vs Laguna XS 2.1 Kimi K2.7 vs Qwythos 9B Kimi K2.7 vs LongCat-2.0 Kimi K2.7 vs Hy3 Kimi K2.7 vs Gemma-4 12B CoderRead more on agentos.guide: /kimi-code, /kimi-hermes, /kimi-modes-head-to-head, /three-dragons
Kimi K2.7 — frequently asked
What is Kimi K2.7?
Kimi K2.7 is Moonshot AI's AI model — The heavy lifter — frontier coder at flat-rate. It has a 256K tokens context window and was released 2026-06.
How good is Kimi K2.7 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 7.46/10 across 25 scored tasks, with 1 gold, 2 silver and 0 bronze medals.
How much does Kimi K2.7 cost?
Flat plan (no per-token bill). Available on Moonshot's flat-rate subscription plan — no per-token billing for individual builders. The plan covers all three speed modes (Fast, No-Think, Quality). Vendor: Moonshot AI (moonshot.ai), based in Beijing.
Where can I see Kimi K2.7 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.