Get the Agent OS + join 4,000+ founders inside the AI Profit Boardroom → Join AIPB ($59/mo)

GoldieBench deep dive

MiMo-V2.6 Pro's full benchmark breakdown.

Open weights that score level with Opus 5 on agents, for cents. Every number below is either one of my own judged one-shot builds - playable on this site - or an outside result with its source linked. Nothing display-only, nothing vibes.

Data refreshed
2026-09-22
Scored tasks
50
Reading time
10 min
External sources
4

01 · The headline numbers

GoldieBench average
6.35/10
50 scored one-shot tasks
Board rank
#23
of 25 ranked frontier models
Task medals
5🥇 8🥈 4🥉
outright wins on shared briefs
Context
1,000,000 tokens
$0.435 in / $0.87 out per M tokens

02 · Every benchmark, bar by bar

All 50 scored tasks, best first. Hover for the judge's comment; click through for every model's build on that task. Or overlay other models in the interactive graphs.

03 · Where it wins - with the judge's own words

“Stunning null-geodesic ray-traced render with convincing photon ring, Doppler-beamed disk warping over/under the shadow, and a swirling lensed starfield — the physics HUD (photon sphere, ISCO, critical curve) is grounded and the whole thing reads at 60fps. Ties the field best; only nitpick is minor label overlap near the shadow edge.”

- the judge's verdict, unedited

“Gorgeous raymarched metaball lava lamp with a proper glass vessel, cap, base, warm bulb glow and molten blobs at varied heights — clearly reads as an authentic lava lamp; the classic silhouette, glass refraction and polished typography make this a top-tier entry.”

- the judge's verdict, unedited

“Gorgeous authentic Matrix rain with multi-layer depth, katakana glyphs, glowing bright heads, and a visible shockwave ripple; polished HUD, palette/flow controls, boot text and scanlines make it a task winner. Only nitpick: the glitch title overlays the subtitle slightly, but visually it nails the brief.”

- the judge's verdict, unedited

“Renders a gorgeous 3D synthwave arena with detailed ship, glowing perspective grid, floating debris, and a fully polished HUD (hull/shield/boost meters, radar, boss bar, combo) backed by scheduled synth music and juicy SFX. Strong on-brief execution that rivals the field best; only minor concern is the sheer scope of a single-file 3D build, but the visuals and systems clearly deliver.”

- the judge's verdict, unedited

“Gorgeous rendered scene—gradient cyberpunk sky, dense emissive-window skyscrapers, overpass structures, neon road glow and a detailed player craft with full combat HUD, minimap, bolts and 'RAM IMPACT' feedback all working; the polished driving-shooter loop clearly beats the generic entries, with only minor HUD text legibility as a small weakness.”

- the judge's verdict, unedited

04 · Where it struggles - quoted, not hidden

Boards that hide the weak rows are brochures. These are MiMo-V2.6 Pro's lowest scored builds, verdicts unedited:

“The screenshot is entirely black with no visible solar system, HUD, or UI panels — despite detailed orbital-mechanics source code, nothing rendered on screen. A non-rendering build scores in the failure range regardless of code quality.”

- the judge's verdict, unedited

“The screenshot shows only a 'RIFT ENGINE FAULT' error message on a dark background — the game caught an exception during init and never rendered, so none of the ambitious sprites, combat, or inventory HUD are visible. Non-functional build regardless of source polish.”

- the judge's verdict, unedited

“The HUD renders nicely (health/ammo/armor/wanted stars/minimap/controls), but the 3D scene is entirely black — no city, streets, pedestrians, or player visible, so the core on-foot GTA gameplay does not appear to render. The 'NO COVER NEAR' banner suggests scripts partially run, but the world is missing.”

- the judge's verdict, unedited

“Polished HUD (quest chip, HP/ST/XP meters, minimap, control hints) renders cleanly, but the 3D world itself is entirely absent — no terrain, sky, characters or enemies appear, leaving a black void where the RPG scene should be. Likely a JS error (note the suspicious 'MeshStandardStandard' typo) halting the render, making this non-functional despite promising source.”

- the judge's verdict, unedited

05 · Category breakdown vs the whole field

Games (23 tasks)

MiMo-V2.6 Pro5.6
field avg7.1

Others (3 tasks)

MiMo-V2.6 Pro6.7
field avg7.4

Pages (3 tasks)

MiMo-V2.6 Pro8.2
field avg7.7

Sims (12 tasks)

MiMo-V2.6 Pro6.6
field avg6.9

Visuals (9 tasks)

MiMo-V2.6 Pro7.2
field avg6.9

06 · Outside signals - every row sourced

My bench measures one thing: judged one-shot builds. These outside rows measure other things - kept separate, never blended into the GoldieBench average, every value linked to where it comes from.

From the source guides

MeasureResultSource guide
Artificial Analysis Intelligence Index v4.346.32/mimo.xiaomi.com
DeepSWE v1.171.9%/huggingface.co/XiaomiMiMo
AutomationBench v1.0.653.1%/huggingface.co/XiaomiMiMo
OSWorld-Verified82.0%/huggingface.co/XiaomiMiMo

07 · Head-to-head records

08 · Methodology + honest limits

Every GoldieBench score is a one-shot, single-file build from an identical prompt - no retries, no hand-fixing - rendered for real, screenshotted, and scored 0-10 by one judge model on one rubric across the whole field. Failures score as failures. What this bench does NOT measure: multi-turn agent work, long-context recall, or API latency - that is what the sourced outside rows are for. Full method: /methodology.

Frequently asked questions

01How does MiMo-V2.6 Pro perform on GoldieBench?

MiMo-V2.6 Pro averages 6.35/10 across 50 scored one-shot build tasks, ranking #23 of 25 ranked frontier models, with 5 task golds, 8 silvers and 4 bronzes. Every score is a real judged build you can open and play on this site.

02What is MiMo-V2.6 Pro best at?

Its strongest scored build is Blackhole at 9.0/10. By category it averages Game 5.6, Other 6.7, Page 8.2, Sim 6.6, Visual 7.2 on the bench.

03How much does MiMo-V2.6 Pro cost?

Xiaomi's September 2026 open-weight flagship (MIT licence, 1.02T total / 42B active parameters). $0.435 per million input tokens and $0.87 per million output on OpenRouter, with cache hits at a fraction of a cent, which is roughly a quarter of Grok 4.7 and a twentieth of the closed frontier models it scores level with on agent benchmarks.

Source ledger

  1. 01Official vendor site: mimo.xiaomi.com/mimo-v2-6mimo.xiaomi.com
  2. 02Artificial Analysis Intelligence Index v4.3agentos.guide
  3. 03DeepSWE v1.1agentos.guide
  4. 04Julian's guide: /xiaomi-mimo-v2-6agentos.guide

Every model's breakdown

The same stack Julian uses

Run this stack yourself.

Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.

4,000+founders
258documented wins
38countries
$59/momonthly