Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)
Kilo

Kilo Code

Fable 5-class intelligence at ~59% less. The split-the-cost play.

ContextVaries — Kilo splits planning from execution across multiple models
Pricing~59% less than Fable 5 solo
Tasks tested0
Avg scorecurrently unranked
Medals🥇0 🥈0 🥉0
Release2026-06-16
Official sitekilocode.ai ↗
Official vendor source
Kilo Code is built by Kilo — see the vendor's own product page, pricing, and docs at kilocode.ai.
Visit kilocode.ai →

Reference benchmarks for Kilo Code

These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for Kilo Code is honest about what's measured.

Cheaper-plan quality (GPT-5.5 cohort)
8.3/10
source: /kilo-split
Cost reduction vs solo Fable 5
~59% less
source: /kilo-split

What is Kilo Code?

Kilo Code is the Kilo frontier model with a Varies — Kilo splits planning from execution across multiple models context window, released 2026-06-16. Tagline: Fable 5-class intelligence at ~59% less. The split-the-cost play.. Official source: kilocode.ai.

Pricing detail. Kilo Code is a routing layer that splits planning (heavy model) from execution (cheaper model) so you get Fable-5-class plans driving GPT-5.5-class builds. Total spend lands at ~59% less than running Fable 5 end-to-end.

How I use it inside the Agent OS. Used inside Agent OS as a routing layer: Fable 5 generates the plan, cheaper models execute. Bench scoring pending a head-to-head comparison.

What I built with Kilo Code

Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Kilo Code shipped on the bench: 0 one-shot demos across Varies — Kilo splits planning from execution across multiple models of context. Of those, 0 are scored against the field with my honest 0–10 from the source guides at agentos.guide.

Strengths

  • Kilo's own rubric: Fable 5 plan = 9.1/10, GPT-5.5 plan = 8.3/10 — Kilo isolates where the intelligence actually lives
  • Plan quality stays high while execution cost drops
  • Drop-in for Agent OS — Kilo Split framework already wired

Trade-offs

  • Adds routing complexity — two model providers in one workflow
  • No per-task goldiebench head-to-heads yet

Best for

  • Cost-conscious operators who run high-volume agent loops
  • Multi-step workflows where the plan is the expensive part
  • Teams already paying for Fable 5 who want to keep the plan but drop the execution bill

Every demo by Kilo Code

0 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.

every demo, in a grid · click any one to play

Compare Kilo Code against every other model

Every head-to-head featuring Kilo Code. Verdicts shown for scored pairs.

Kilo Code vs Fusion
Reference-only
Kilo Code vs Claude Opus 5
Reference-only
Kilo Code vs Hermes MoA
Reference-only
Kilo Code vs GPT-5.6 Sol
Reference-only
Kilo Code vs Claude Fable 5
Reference-only
Kilo Code vs Qwen 3.8
Reference-only
Kilo Code vs Grok
Reference-only
Kilo Code vs MiniMax M3
Reference-only
Kilo Code vs Fugu Ultra
Reference-only
Kilo Code vs Kimi K3
Reference-only
Kilo Code vs GLM-5.2
Reference-only
Kilo Code vs Fugu Mini
Reference-only
Kilo Code vs Opus 4.8
Reference-only
Kilo Code vs Kimi K2.7
Reference-only
Kilo Code vs Qwable 5 27B Coder
Reference-only
Kilo Code vs Gemini 3.6 Flash
Reference-only
Kilo Code vs Claude Sonnet 5
Reference-only
Kilo Code vs Qwen 3.7
Reference-only
Kilo Code vs Fugu Ultra 1.1
Reference-only
Kilo Code vs Inkling
Reference-only
Kilo Code vs Agents-A1
Reference-only
Kilo Code vs Gemma 4 12B · MLX
Reference-only
Kilo Code vs Laguna XS 2.1
Reference-only
Kilo Code vs Qwythos 9B
Reference-only
Kilo Code vs LongCat-2.0
Reference-only
Kilo Code vs Hy3
Reference-only
Kilo Code vs Gemma-4 12B Coder
Reference-only
Kilo Code vs DeepSeek V4 Flash
Reference-only
Kilo Code vs DeepSeek V4 Pro
Reference-only
Kilo Code vs Kimi K2.7 · Fast
Reference-only
Kilo Code vs Kimi K2.7 · No-Think
Reference-only
Kilo Code vs Kimi K2.7 · Quality
Reference-only
Kilo Code vs Ornith 1.0
Reference-only
Kilo Code vs Claude Mythos 5
Reference-only

See all 66 comparisons across every model →

Quick pill index

Direct comparisons against every other scored model on the bench:

Kilo Code vs Fusion Kilo Code vs Claude Opus 5 Kilo Code vs Hermes MoA Kilo Code vs GPT-5.6 Sol Kilo Code vs Claude Fable 5 Kilo Code vs Qwen 3.8 Kilo Code vs Grok Kilo Code vs MiniMax M3 Kilo Code vs Fugu Ultra Kilo Code vs Kimi K3 Kilo Code vs GLM-5.2 Kilo Code vs Fugu Mini Kilo Code vs Opus 4.8 Kilo Code vs Kimi K2.7 Kilo Code vs Qwable 5 27B Coder Kilo Code vs Gemini 3.6 Flash Kilo Code vs Claude Sonnet 5 Kilo Code vs Qwen 3.7 Kilo Code vs Fugu Ultra 1.1 Kilo Code vs Inkling Kilo Code vs Agents-A1 Kilo Code vs Gemma 4 12B · MLX Kilo Code vs Laguna XS 2.1 Kilo Code vs Qwythos 9B Kilo Code vs LongCat-2.0 Kilo Code vs Hy3 Kilo Code vs Gemma-4 12B Coder

Read more on agentos.guide: /kilo-split

Kilo Code — frequently asked

What is Kilo Code?

Kilo Code is Kilo's AI model — Fable 5-class intelligence at ~59% less. The split-the-cost play. It has a Varies (Kilo dispatches across models) context window and was released 2026-06-16.

How good is Kilo Code at coding and one-shot builds?

It has 0 live demos on GoldieBench but no curated 0-10 verdicts yet — it is unranked until scored.

How much does Kilo Code cost?

~59% less than Fable 5 solo. Kilo Code is a routing layer that splits planning (heavy model) from execution (cheaper model) so you get Fable-5-class plans driving GPT-5.5-class builds. Total spend lands at ~59% less than running Fable 5 end-to-end.

Where can I see Kilo Code demos?

Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly