Skip to content

Method

How we grade

Every number on this site comes from 74 models, 194 devices and 17 labs, each figure with its public source, and a speed model checked against 193 published measurements.

Detection

Your browser reports the name of its graphics chip through WebGPU or WebGL. We match it against our device list; when a card is sold with several memory sizes, or Safari hides the chip, we ask you to pick. Nothing about your hardware is sent anywhere: the grading runs on your device.

Memory

A model needs room for its weights, for the cache that holds the conversation, and for the runtime's working buffers. We use real GGUF file sizes when a public repository has them and the bits per weight from llama.cpp otherwise. The cache grows with context, which is why the context selector changes the grades.

memory = weights + KV cache + runtime
weights  = parameters × bits per weight / 8      (Q4_K_M ≈ 4.89 bits; real GGUF size when known)
KV cache = 2 × layers × KV heads × head size × 2 bytes × context
           (windowed layers stop at their window)

Where the memory lives

On a graphics card, about 1 GB is kept by the driver and the desktop. On a Mac, macOS lets the GPU use about two thirds of memory up to 32 GB and three quarters above. Several cards share the model in proportion to their memory, and whatever doesn't fit spills to system RAM, which is far slower. Generation speed is limited by how fast memory can be read, so we estimate it from bandwidth plus a fixed cost per token.

seconds per token = bytes read per token / (efficiency × bandwidth) + fixed cost
bytes read        = active weights + KV cache in use

How far off we are

We fit the speed model on published llama.cpp benchmarks and report the error by predicting each measurement without using it. The figure below is the median difference between our estimate and the real speed.

BackendMeasurementsBandwidth usedMedian error
CUDA3366%12%
Metal6693%4.9%
ROCm2479%21%
Vulkan3485%18.3%

Grades

Grades follow speed, anchored on reading: people read about 5 to 8 tokens a second, so below that the model is slower than you read. A tight fit can't go above B and spilling to system RAM can't go above C, because both break down in long conversations.

  • SExcellent≥ 40 tok/s
  • AGood≥ 20 tok/s
  • BUsable≥ 10 tok/s · cap if tight
  • CSlow≥ 5 tok/s · cap if spilling
  • DBarely runs< 5 tok/s
  • FDoesn't fitdoesn't fit

What we don't know

Speeds are for one user generating text with llama.cpp-style runtimes. Prompt reading, other runtimes, drivers and thermals change real numbers. Some laptop chips vary by maker, and browsers can't report system RAM, so check the values in the hardware bar.