MiniMax M3
1M-context frontier model at $0.30/M tokens — cheapest big-context model on the bench.
Reference benchmarks for MiniMax M3
These are external benchmarks I pulled from the source comparison guides on agentos.guide — SWE-bench Verified, DRACO, Kilo plan rubric, build-time measurements, vendor-reported coding scores. They are not goldiebench medal scores (those come only from same-prompt one-shot creative coding tasks in the matrix). I surface them here so the spec sheet for MiniMax M3 is honest about what's measured.
What is MiniMax M3?
MiniMax M3 is the MiniMax frontier model with a 1,048,576-token context — matches GLM-5.2 and Fable 5 context window, released 2026-06-18. Tagline: 1M-context frontier model at $0.30/M tokens — cheapest big-context model on the bench.. Official source: minimax.io.
Pricing detail. MiniMax M3 is the cheapest 1M-context frontier model on the bench — roughly 1/200th the per-call cost of OpenRouter Fusion and 1/30th of Claude Opus 4.8. Designed for high-volume agent workloads where context length matters but per-call budget is tight.
How I use it inside the Agent OS. Bench prompts dispatched via OpenRouter. Scored by Claude judge against the same 42 prompts every other model ran.
What I built with MiniMax M3
Every model on Goldie Bench gets the same fixed prompt set — one shot, single HTML file out — and I score the result 0–10 inside the Agent Operating System. Here's what MiniMax M3 shipped on the bench: 47 one-shot demos across 1,048,576-token context — matches GLM-5.2 and Fable 5 of context. Of those, 47 are scored against the field with my honest 0–10 from the source guides at agentos.guide.
Strengths
- 1M token context — full repo / full deep-research corpus fits in one call
- $0.30/M input is roughly 1/30th of Opus 4.8 — built for high-volume agent loops
- Solid one-shot HTML output — clean structure on game and visual prompts
Trade-offs
- Less polished than Fusion's panel-ensembled output on the toughest deep builds
- Newer model — less community calibration vs Fable 5 / Opus 4.8
Best for
- High-volume agent workflows where per-call cost dominates
- 1M-context tasks (whole-repo refactors, deep-research synthesis)
- Drop-in cheaper alternative to GLM-5.2 with comparable 1M context
Every benchmark — MiniMax M3's full scorecard
All 47 scored tasks, best first — the judge's 0–10 on the same rubric as the whole field. Click any bar for that task's cross-model page, or open this scorecard in the interactive graphs. Full editorial breakdown with judge quotes and sourced outside research: the MiniMax M3 deep dive →.
Every demo by MiniMax M3
47 live demos, sorted by category. Click any tile to play the actual one-shot result. Verdicts and 0–10 scores are pulled from the source guides where I posted them publicly.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVECompare MiniMax M3 against every other model
Every head-to-head featuring MiniMax M3. Verdicts shown for scored pairs.
See all 66 comparisons across every model →
Quick pill index
Direct comparisons against every other scored model on the bench:
MiniMax M3 vs Fusion MiniMax M3 vs Claude Opus 5 MiniMax M3 vs Hermes MoA MiniMax M3 vs GPT-5.6 Sol MiniMax M3 vs Claude Fable 5 MiniMax M3 vs Qwen 3.8 MiniMax M3 vs Grok MiniMax M3 vs Fugu Ultra MiniMax M3 vs Kimi K3 MiniMax M3 vs GLM-5.2 MiniMax M3 vs Fugu Mini MiniMax M3 vs Muse Spark 1.2 MiniMax M3 vs Opus 4.8 MiniMax M3 vs Kimi K2.7 MiniMax M3 vs Qwable 5 27B Coder MiniMax M3 vs Gemini 3.6 Flash MiniMax M3 vs Claude Sonnet 5 MiniMax M3 vs Qwen 3.7 MiniMax M3 vs Fugu Ultra 1.1 MiniMax M3 vs Inkling MiniMax M3 vs Agents-A1 MiniMax M3 vs Gemma 4 12B · MLX MiniMax M3 vs Laguna XS 2.1 MiniMax M3 vs Qwythos 9B MiniMax M3 vs LongCat-2.0 MiniMax M3 vs Hy3 MiniMax M3 vs Gemma-4 12B CoderRead more on agentos.guide:
MiniMax M3 — frequently asked
What is MiniMax M3?
MiniMax M3 is MiniMax's AI model — 1M-context frontier model at $0.30/M tokens — cheapest big-context model on the bench. It has a 1M tokens context window and was released 2026-06-18.
How good is MiniMax M3 at coding and one-shot builds?
On the GoldieBench one-shot build benchmark it averages 7.97/10 across 47 scored tasks, with 2 gold, 1 silver and 4 bronze medals.
How much does MiniMax M3 cost?
$0.30 / 1M input tokens, $1.50 / 1M output. MiniMax M3 is the cheapest 1M-context frontier model on the bench — roughly 1/200th the per-call cost of OpenRouter Fusion and 1/30th of Claude Opus 4.8. Designed for high-volume agent workloads where context length matt
Where can I see MiniMax M3 demos?
Every one-shot build is live and playable on this page and on the GoldieBench compare matrix — same prompt as every other model, no retries.
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.