Throughput is modelled from each chip's unified-memory bandwidth and GPU FLOPS (decode is bandwidth-bound, prefill is compute-bound). The "your browser" figure is measured live.
Drag to orbit · Scroll/pinch to zoom · Click a bar to re-run that chip · R: run all · M: next model