Qwythos 9B
A Claude-style creative & reasoning 9B with a full 1M-token context — the local writer & thinker.
What is Qwythos 9B?
Qwythos 9B is the Richard Young · DeepNeuro (abliterated build of empero-ai's Qwythos, Qwen3.5 base) frontier model with a 1,048,576 tokens (YaRN, Qwen3.5-9B base) context window, released 2026-06. Tagline: A Claude-style creative & reasoning 9B with a full 1M-token context — the local writer & thinker.. Official source: ollama.com/richardyoung/qwythos-9b-abliterated.
Pricing detail. Qwythos (model ID: richardyoung/qwythos-9b-abliterated) is an abliterated build of empero-ai's Qwythos-9B-Claude-Mythos — a Claude-style creative & reasoning model on a Qwen3.5-9B base, post-trained on Claude Mythos & Fable traces, with a 1M-token context. Thinking model (<think>), native function-calling. Refusals trimmed via the Heretic library (53/100, KL 0.0066). Q4_K_M is 5.6GB, runs 100% offline.
How I use it inside the Agent OS. Loaded in Ollama as the local long-context + creative option. Now fully bench-scored (42/42, avg 2.98): it handles simpler page and arcade builds, but most of its ambitious one-shot 3D/WebGL attempts errored out — an honest read on a free 9B's ceiling for hard one-shots.
What I built with Qwythos 9B
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what Qwythos 9B shipped on the bench: 42 one-shot demos across 1,048,576 tokens (YaRN, Qwen3.5-9B base) of context. Of those, 42 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- Full 1M-token context locally — matches frontier-cloud models on context length, at $0
- Fast on a Mac — a steady ~52 tokens/sec on real builds (M4 Max), just 5.6GB on disk
- Open-weights (Qwen3.5 base), runs 100% on-device via Ollama — free and private
- Best bench builds: Nordic Crypt 6.8, Racing 6.5, Neon Blaster 6.4, Arcade 6.2, Landing 6.0
Trade-offs
- 9B ceiling — bench-scored 2.98/10 avg (42/42), last on the one-shot creative-coding suite
- Ambitious 3D/WebGL one-shots often ship with real bugs — bare `three` imports, const reassignment, undefined refs — so they render blank
- Strong on simpler page/arcade/2D tasks; falls apart on complex graphics where frontier models hold up
Best for
- Long-context local tasks (large file refactors, multi-file analysis) where 1M ctx matters
- Cost-zero daily coding work on a consumer Mac
- Workflows where you want an unfiltered local coder without API guardrails
Every benchmark — Qwythos 9B's full scorecard
All 42 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the Qwythos 9B deep dive →.
Every demo by Qwythos 9B
42 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare Qwythos 9B against every other model
Every head-to-head featuring Qwythos 9B. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
Qwythos 9B vs Fusion Qwythos 9B vs Claude Opus 5 Qwythos 9B vs Hermes MoA Qwythos 9B vs GPT-5.6 Sol Qwythos 9B vs Claude Fable 5 Qwythos 9B vs Qwen 3.8 Qwythos 9B vs Grok Qwythos 9B vs MiniMax M3 Qwythos 9B vs Fugu Ultra Qwythos 9B vs Kimi K3 Qwythos 9B vs GLM-5.2 Qwythos 9B vs Fugu Mini Qwythos 9B vs Claude Opus 5.5 Qwythos 9B vs Muse Spark 1.2 Qwythos 9B vs Opus 4.8 Qwythos 9B vs Kimi K2.7 Qwythos 9B vs Grok 4.7 Qwythos 9B vs Qwable 5 27B Coder Qwythos 9B vs Gemini 3.6 Flash Qwythos 9B vs Claude Sonnet 5 Qwythos 9B vs Qwen 3.7 Qwythos 9B vs Fugu Ultra 1.1 Qwythos 9B vs MiMo-V2.6 Pro Qwythos 9B vs Inkling Qwythos 9B vs Grok 4.6 Qwythos 9B vs Agents-A1 Qwythos 9B vs Gemma 4 12B · MLX Qwythos 9B vs Laguna XS 2.1 Qwythos 9B vs LongCat-2.0 Qwythos 9B vs Hy3 Qwythos 9B vs Gemma-4 12B CoderRead more on agentos.guide:
Qwythos 9B — frequently asked
What is Qwythos 9B?
Qwythos 9B is Richard Young · DeepNeuro (abliterated build of empero-ai's Qwythos, Qwen3.5 base)'s AI model — A Claude-style creative & reasoning 9B with a full 1M-token context — the local writer & thinker. It has a 1M tokens context window and was released 2026-06.
How good is Qwythos 9B at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 2.98/10 across 42 scored tasks, with 0 gold, 0 silver and 0 bronze medals.
How much does Qwythos 9B cost?
Free · runs locally. Qwythos (model ID: richardyoung/qwythos-9b-abliterated) is an abliterated build of empero-ai's Qwythos-9B-Claude-Mythos — a Claude-style creative & reasoning model on a Qwen3.5-9B base, post-trained on Claude Mythos & Fa
Where can I see Qwythos 9B demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.