Method
How we grade
Every number on this site comes from 74 models, 194 devices and 17 labs, each figure with its public source, and a speed model checked against 193 published measurements.
Detection
Your browser reports the name of its graphics chip through WebGPU or WebGL. We match it against our device list; when a card is sold with several memory sizes, or Safari hides the chip, we ask you to pick. Nothing about your hardware is sent anywhere: the grading runs on your device.
Memory
A model needs room for its weights, for the cache that holds the conversation, and for the runtime's working buffers. We use real GGUF file sizes when a public repository has them and the bits per weight from llama.cpp otherwise. The cache grows with context, which is why the context selector changes the grades.
memory = weights + KV cache + runtime
weights = parameters × bits per weight / 8 (Q4_K_M ≈ 4.89 bits; real GGUF size when known)
KV cache = 2 × layers × KV heads × head size × 2 bytes × context
(windowed layers stop at their window)Where the memory lives
On a graphics card, about 1 GB is kept by the driver and the desktop. On a Mac, macOS lets the GPU use about two thirds of memory up to 32 GB and three quarters above. Several cards share the model in proportion to their memory, and whatever doesn't fit spills to system RAM, which is far slower. Generation speed is limited by how fast memory can be read, so we estimate it from bandwidth plus a fixed cost per token.
seconds per token = bytes read per token / (efficiency × bandwidth) + fixed cost bytes read = active weights + KV cache in use
How far off we are
We fit the speed model on published llama.cpp benchmarks and report the error by predicting each measurement without using it. The figure below is the median difference between our estimate and the real speed.
| Backend | Measurements | Bandwidth used | Median error |
|---|---|---|---|
| CUDA | 33 | 66% | 12% |
| Metal | 66 | 93% | 4.9% |
| ROCm | 24 | 79% | 21% |
| Vulkan | 34 | 85% | 18.3% |
Grades
Grades follow speed, anchored on reading: people read about 5 to 8 tokens a second, so below that the model is slower than you read. A tight fit can't go above B and spilling to system RAM can't go above C, because both break down in long conversations.
- SExcellent≥ 40 tok/s
- AGood≥ 20 tok/s
- BUsable≥ 10 tok/s · cap if tight
- CSlow≥ 5 tok/s · cap if spilling
- DBarely runs< 5 tok/s
- FDoesn't fitdoesn't fit
What we don't know
Speeds are for one user generating text with llama.cpp-style runtimes. Prompt reading, other runtimes, drivers and thermals change real numbers. Some laptop chips vary by maker, and browsers can't report system RAM, so check the values in the hardware bar.