Mlx Speedtest
Mlx Speedtest — auto-discovered task.
What I asked each model — the Mlx Speedtest prompt
Every model on this page got this exact prompt inside the Agent Operating System: Mlx Speedtest — auto-discovered task.
Single HTML file out. No iteration. No examples in the system prompt. Whatever each model produced on the first run is what's on this page. 9 frontier models have attempted it so far: Gemini 3.6 Flash, GPT-5.6 Sol, Inkling, Kimi K3, Muse Spark 1.2, Claude Opus 5, Qwen 3.8, DeepSeek V4 Pro, DeepSeek V4 Flash.
Why this task matters. Mlx Speedtest is a textbook test of other-class capability — the kind of build that exposes whether a model is doing pattern-matching or actual reasoning. Shipping this cleanly is the floor for what I expect from a frontier model — every model on the leaderboard should at least attempt it.
How each model handled Mlx Speedtest
Ranked by my 0–10 score from the source comparison guides on agentos.guide. Click any to play the actual one-shot HTML the model produced.
What I saw: Strong, polished MLX benchmark dashboard with a 3D Three.js tensor core, orbital rings, glassmorphic HUD, chip presets and clean metric cards — clearly shippable. Falls just short of the field's best since the screenshot shows an idle state (0 GB/s, 0.0 TFLOPS, 1.0 tok/s) without visible running/animated benchmark data to prove the speedtest actually sweeps.
What I saw: Strong, cohesive dark dashboard with a beautifully rendered speedometer gauge, sensible MLX-flavored task metadata (Llama 3.2 3B, 4-bit, Metal backend) and clean live-sample layout that clearly beats the generic field. Minor weakness: captured in a paused/pre-run state so metrics read empty (0.0, blank ms/GB/tok/W), leaving the payoff unverified in the shot.
What I saw: Strong atom/orbit-ring 3D visual with a polished shimmer title and clean readout typography, and it clearly renders on-brand. However the 'speedtest' is purely cosmetic random noise with no real benchmark logic, and the big readout awkwardly overlaps the core sphere, keeping it just shy of the field's best.
What I saw: Strong 3D-backed speedtest with polished neon gauge, live stats, and orbiting torus knot that renders cleanly and reads as on-brief; weakened by the busy wireframe knot obscuring the central gauge/readout and the sparkline area appearing as an empty blurred bar, reducing legibility.
What I saw: Strong polished dashboard shell with clean panels, live metrics, sparklines and a benchmarking task list that reads convincingly on-brief. Weakened by the empty 3D center canvas (just faint ring, no visible model/animation) and blank throughput timeline chart, plus raw unformatted tok/s numbers, which keep it below the field's best.
What I saw: Strong visual polish with a real 3D A·B=C matrix visualization, four genuine GEMM kernels, and a clean HUD showing matrix size/passes/verification — but the throughput readouts are all stuck at 0.00 GFLOP/s with empty bars, so the core 'speedtest' benchmark isn't actually surfacing results, undercutting the brief.
What I saw: Strong on-brief render: crisp MLX branding, 6 discovered models with tok/s estimates, phased benchmark steps, a 3D gauge with needle, live log, platform specs and a throughput chart — highly polished and cohesive. Slight weakness is the paused/loading state means the gauge reads 0 and results table is empty in this shot, but the simulation architecture is clearly complete and ships above the field.
Demo on the bench. Not scored yet — play it and form your own opinion.
Demo on the bench. Not scored yet — play it and form your own opinion.
The winner on Mlx Speedtest
Qwen 3.8 took gold on this task. Polished MLX dashboard.
What I saw: Strong on-brief render: crisp MLX branding, 6 discovered models with tok/s estimates, phased benchmark steps, a 3D gauge with needle, live log, platform specs and a throughput chart — highly polished and cohesive. Slight weakness is the paused/loading state means the gauge reads 0 and results table is empty in this shot, but the simulation architecture is clearly complete and ships above the field.
See Qwen 3.8's full model card: /models/qoder. Direct head-to-head against the runner-up: Qwen 3.8 vs GPT-5.6 Sol.
Every attempt — live, playable
Side by side. Click any tile to run that model's actual one-shot HTML in a new tab.
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVE
▶ LIVEHow I scored Mlx Speedtest — methodology
Three axes, 0–10 each, averaged. Runs: drop the .html in a browser; if it opens to a broken page, it scores zero. Hits the brief: did the model ship the thing the prompt asked for, or a different thing it found easier. Looks good: visual polish, motion, interactivity — where most of the gap between gold and silver lives.
My scores trace back to the source comparison guides on agentos.guide. See the full methodology page for data provenance, including which source guide each cell's score came from.
Related
More other benchmarks: all tasks in the Other category · See the best AI model for Mlx Speedtest · Back to the leaderboard
Run this stack yourself.
Every demo on this bench was built inside the Agent Operating System — one prompt, one shot, single HTML file out. The Agent OS, the prompts, the templates, the weekly walkthroughs and 4,000+ founders shipping with it every day all live inside the AI Profit Boardroom.