Kilo Code
Fable 5-class intelligence at ~59% less. The split-the-cost play.
Reference benchmarks for Kilo Code
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Kilo Code is honest about what's measured.
What is Kilo Code?
Kilo Code is the Kilo frontier model with a Varies — Kilo splits planning from execution across multiple models context window, released 2026-06-16. Tagline: Fable 5-class intelligence at ~59% less. The split-the-cost play.. Official source: kilocode.ai.
Pricing detail. Kilo Code is a routing layer that splits planning (heavy model) from execution (cheaper model) so you get Fable-5-class plans driving GPT-5.5-class builds. Total spend lands at ~59% less than running Fable 5 end-to-end.
How I use it inside the Agent OS. Used inside Agent OS as a routing layer: Fable 5 generates the plan, cheaper models execute. Bench scoring pending a head-to-head comparison.
What I built with Kilo Code
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Kilo Code shipped on the bench: 0 one-shot demos across Varies — Kilo splits planning from execution across multiple models of context. Of those, 0 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Kilo's own rubric: Fable 5 plan = 9.1/10, GPT-5.5 plan = 8.3/10 — Kilo isolates where the intelligence actually lives
- Plan quality stays high while execution cost drops
- Drop-in for Agent OS — Kilo Split framework already wired
Trade-offs
- Adds routing complexity — two model providers in one workflow
- No per-task goldiebench head-to-heads yet
Best for
- Cost-conscious operators who run high-volume agent loops
- Multi-step workflows where the plan is the expensive part
- Teams already paying for Fable 5 who want to keep the plan but drop the execution bill
Every demo by Kilo Code
0 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
Compare Kilo Code against every other model
Every head-to-head featuring Kilo Code. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Kilo Code vs Fusion Kilo Code vs Claude Opus 5 Kilo Code vs Hermes MoA Kilo Code vs GPT-5.6 Sol Kilo Code vs Claude Fable 5 Kilo Code vs Qwen 3.8 Kilo Code vs Grok Kilo Code vs MiniMax M3 Kilo Code vs Fugu Ultra Kilo Code vs Kimi K3 Kilo Code vs GLM-5.2 Kilo Code vs Fugu Mini Kilo Code vs Opus 4.8 Kilo Code vs Kimi K2.7 Kilo Code vs Qwable 5 27B Coder Kilo Code vs Gemini 3.6 Flash Kilo Code vs Claude Sonnet 5 Kilo Code vs Qwen 3.7 Kilo Code vs Fugu Ultra 1.1 Kilo Code vs Inkling Kilo Code vs Agents-A1 Kilo Code vs Gemma 4 12B · MLX Kilo Code vs Laguna XS 2.1 Kilo Code vs Qwythos 9B Kilo Code vs LongCat-2.0 Kilo Code vs Hy3 Kilo Code vs Gemma-4 12B CoderRead more on agentos.guide: /kilo-split
Kilo Code — frequently asked
What is Kilo Code?
Kilo Code is Kilo's AI model — Fable 5-class intelligence at ~59% less. The split-the-cost play. It has a Varies (Kilo dispatches across models) context window and was released 2026-06-16.
How good is Kilo Code at coding and one-shot builds?
It has 0 live demos on GoldieBench but no curated 0-10 verdicts yet — it is unranked until scored.
How much does Kilo Code cost?
~59% less than Fable 5 solo. Kilo Code is a routing layer that splits planning (heavy model) from execution (cheaper model) so you get Fable-5-class plans driving GPT-5.5-class builds. Total spend lands at ~59% less than running Fable 5 end-to-end.
Where can I see Kilo Code demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.