GoldieBench deep dive
MiMo-V2.6 Pro's full benchmark breakdown.
Open weights that score level with Opus 5 on agents, for cents. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.
01 · The headline numbers
02 · Every benchmark, bar by bar
All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.
03 · Where it wins - with the judge's own words
“Stunning null-geodesic ray-traced render with convincing photon ring, Doppler-beamed disk warping over/under the shadow, and a swirling lensed starfield — the physics HUD (photon sphere, ISCO, critical curve) is grounded and the whole thing reads at 60fps. Ties the field best; only nitpick is minor label overlap near the shadow edge.”
“Gorgeous raymarched metaball lava lamp with a proper glass vessel, cap, base, warm bulb glow and molten blobs at varied heights — clearly reads as an authentic lava lamp; the classic silhouette, glass refraction and polished typography make this a top-tier entry.”
“Gorgeous authentic Matrix rain with multi-layer depth, katakana glyphs, glowing bright heads, and a visible shockwave ripple; polished HUD, palette/flow controls, boot text and scanlines make it a task winner. Only nitpick: the glitch title overlays the subtitle slightly, but visually it nails the brief.”
“Renders a gorgeous 3D synthwave arena with detailed ship, glowing perspective grid, floating debris, and a fully polished HUD (hull/shield/boost meters, radar, boss bar, combo) backed by scheduled synth music and juicy SFX. Strong on-brief execution that rivals the field best; only minor concern is the sheer scope of a single-file 3D build, but the visuals and systems clearly deliver.”
“Gorgeous rendered scene—gradient cyberpunk sky, dense emissive-window skyscrapers, overpass structures, neon road glow and a detailed player craft with full combat HUD, minimap, bolts and 'RAM IMPACT' feedback all working; the polished driving-shooter loop clearly beats the generic entries, with only minor HUD text legibility as a small weakness.”
04 · Where it struggles - quoted, not hidden
Boards that hide the weak rows are brochures. These are MiMo-V2.6 Pro's lowest scored builds, verdicts unedited:
“The screenshot is entirely black with no visible solar system, HUD, or UI panels — despite detailed orbital-mechanics source code, nothing rendered on screen. A non-rendering build scores in the failure range regardless of code quality.”
“The screenshot shows only a 'RIFT ENGINE FAULT' error message on a dark background — the game caught an exception during init and never rendered, so none of the ambitious sprites, combat, or inventory HUD are visible. Non-functional build regardless of source polish.”
“The HUD renders nicely (health/ammo/armor/wanted stars/minimap/controls), but the 3D scene is entirely black — no city, streets, pedestrians, or player visible, so the core on-foot GTA gameplay does not appear to render. The 'NO COVER NEAR' banner suggests scripts partially run, but the world is missing.”
“Polished HUD (quest chip, HP/ST/XP meters, minimap, control hints) renders cleanly, but the 3D world itself is entirely absent — no terrain, sky, characters or enemies appear, leaving a black void where the RPG scene should be. Likely a JS error (note the suspicious 'MeshStandardStandard' typo) halting the render, making this non-functional despite promising source.”
05 · Category breakdown vs the whole field
Games (23 tasks)
Others (3 tasks)
Pages (3 tasks)
Sims (12 tasks)
Visuals (9 tasks)
06 · Outside signals - every row sourced
My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.
From the source guides
| Measure | Result | Source guide |
|---|---|---|
| Artificial Analysis Intelligence Index v4.3 | 46.32 | /mimo.xiaomi.com |
| DeepSWE v1.1 | 71.9% | /huggingface.co/XiaomiMiMo |
| AutomationBench v1.0.6 | 53.1% | /huggingface.co/XiaomiMiMo |
| OSWorld-Verified | 82.0% | /huggingface.co/XiaomiMiMo |
07 · Head-to-head records
08 · Methodology + honest limits
Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.
Frequently asked questions
01How does MiMo-V2.6 Pro perform on GoldieBench?
MiMo-V2.6 Pro averages 6.35/10 across 50 scored one-shot build tasks, ranking #23 of 25 ranked frontier models, with 5 task golds, 8 silvers and 4 bronzes. Every score is a real judged build you can open and play on this site.
02What is MiMo-V2.6 Pro best at?
Its strongest scored build is Blackhole at 9.0/10. By category it averages Game 5.6, Other 6.7, Page 8.2, Sim 6.6, Visual 7.2 on the bench.
03How much does MiMo-V2.6 Pro cost?
Xiaomi's September 2026 open-weight flagship (MIT licence, 1.02T total / 42B active parameters). $0.435 per million input tokens and $0.87 per million output on OpenRouter, with cache hits at a fraction of a cent, which is roughly a quarter of Grok 4.7 and a twentieth of the closed frontier models it scores level with on agent benchmarks.
Source ledger
- 01Official vendor site: mimo.xiaomi.com/mimo-v2-6mimo.xiaomi.com
- 02Artificial Analysis Intelligence Index v4.3agentos.guide
- 03DeepSWE v1.1agentos.guide
- 04Julian's guide: /xiaomi-mimo-v2-6agentos.guide
Every model's breakdown
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.