⭐ Built in the Agent OS — get it inside the AI Profit BoardroomJoin AIPB →

MLX Speedtest

Apple-silicon LLM inference race · prompt processing + token generation on MLX
Measuring your browser's fp32 matmul throughput…
IDLE
Generation
— tok/s
Prompt
— tok/s
Time to 1st token
Peak memory
LEADERBOARD · generation tok/s
Throughput is modelled from each chip's unified-memory bandwidth and GPU FLOPS (decode is bandwidth-bound, prefill is compute-bound). The "your browser" figure is measured live.
Drag to orbit · Scroll/pinch to zoom · Click a bar to re-run that chip · R: run all · M: next model