Calibrated on 193 published benchmarks
What AI can your computer run?
We read your GPU or Mac and grade every open model for it: memory it needs, speed you'll get, and the command to run it.
Your hardware
- Memory
- GB
- Bandwidth
- GB/s
- Context
Most intelligent on your …
Intelligence (ECI) of the most intelligent models that load · colour shows how well they run
Expected speed
Tokens per second on your … for the most intelligent models that run well
Models available with more memory
Open models that run well (grade B or better) at each memory size
Scores from Oct 8, 2026 · 0 new models in the last two months
For any hardware
Every open model and device we measure, independent of your machine
Intelligence
The most intelligent open models, ranked by the Epoch Capabilities Index (ECI)
Intelligence
Epoch Capabilities Index (ECI) · Higher is better
Intelligence by memory
Up and to the left is better. The line joins the models that no smaller model beats.
Best for its size
Speed by device
Output tokens per second for one model at Q4_K_M on popular devices · Higher is better
on popular devices
tok/s · Q4_K_M · 8K
Memory and context
Memory the most intelligent models need at Q4_K_M; a longer context needs more
Memory needed with a 8K context
GB at Q4_K_M · Lower is better
Benchmarks by task
Epoch AI's own evaluations, the best setting it ran for each open model
Reasoning
GPQA Diamond · 0 open models measured · Higher is better
Math
OTIS Mock AIME · 0 open models measured · Higher is better
Knowledge
SimpleQA Verified · 0 open models measured · Higher is better
Coding
SWE-bench Verified · 0 open models measured · Higher is better
Devices
Open models that run well (grade B or better) on each device, at its standard memory
Models that run well
Out of 0 open models · Higher is better
Open vs closed over time
Highest intelligence (ECI) reached by open and by closed models, month by month
New models
Open models released in the last two months, by intelligence · 0 not scored yet
New models
ECI · Higher is better
Quantization
Memory needs at each quantization with an 8K context · colour shows the quality kept
GB · 8K
Noticeably worseSome lossGoodNear originalOriginal
Methodology and open data
How each grade, speed and memory figure is estimated, and every number in downloadable form
74 models · 194 devices · 17 labs
Open AI models on your own hardware
Running a model at home keeps your data on your machine and costs nothing per message. Whether it works comes down to two numbers: how much memory the model needs, and how fast your memory can feed it. We compute both for every model and every device, from public specs and calibrated against published measurements. The details are in how we grade.
- 01_
Open the page
On the computer you want to use. No account, no download, nothing installed.
- 02_
Check your hardware
We read your GPU or chip from the browser. If it's wrong or hidden, pick it from the list and set your memory.
- 03_
Read the grades
Every model gets S to F for your machine, with the memory it needs and the speed you'll get.
- 04_
Run it
Open a model to see the best quantization for your hardware and the command to start it.
How much memory do you need?
Computed by our engine for Q4_K_M with an 8K context on a dedicated GPU. Mixture-of-experts models load every expert into memory, even though only a few work on each word.
| Memory | What fits | Example |
|---|---|---|
| 8 GB | 15 models, up to Qwen3.5 9B (9.7B) | RTX 4060 |
| 12 GB | 18 models, up to Phi-4 14B (14.7B) | RTX 3060 12GB |
| 16 GB | 23 models, up to Devstral Small 2 24B (24B) | RTX 4060 Ti 16GB |
| 24 GB | 35 models, up to Qwen3.6 35B A3B (36B) | RTX 4090 |
| 32 GB | 37 models, up to Kimi Linear 48B A3B (49.1B) | RTX 5090 |
| 48 GB | 39 models, up to Qwen3 Coder Next (79.7B) | RTX A6000 |
Popular devices
Labs behind the models
Models worth trying first
A short list to start with. The grades above still decide what fits your machine.
Qwen3.5 4B4.7BSmall, quick and multilingual. A good first model for 8 GB cards and laptops.
gpt-oss 20b3.6B activeOpenAI's open-weight reasoning model. Fast because only a fraction of it works per word.
Qwen3.6 35B A3B3B activeA large mixture-of-experts that runs like a small model when it fits in memory.
Gemma 4 12B12BGoogle's mid-size model with vision. Comfortable on 12–16 GB cards.
Devstral Small 2 24B24BBuilt for coding agents. Needs a 16–24 GB card to breathe.
Llama 3.1 8B Instruct8BThe classic 8B. Runs almost anywhere and every tool supports it.
Common questions
The formulas, the data sources and how far off we are. How we grade
How much VRAM do I need to run AI locally?
8 GB runs good small models (3–9B) at Q4. 16 GB opens up the 12–24B range, and 24 GB or more lets you use 30B models and mid-size mixtures of experts. Longer conversations need extra memory for the context.
Can a Mac run AI models without a graphics card?
Yes. Apple Silicon shares one pool of memory between CPU and GPU, so a Mac with 32 GB or more can hold models that wouldn't fit on most graphics cards. macOS keeps about a third of that memory for itself by default.
What is quantization?
Storing a model's numbers with fewer bits. Q4_K_M uses about 4.9 bits per weight instead of 16, so the model takes roughly a third of the memory with a small loss in quality. Lower than Q4 saves more memory but starts to hurt answers.
How do you know which models fit my hardware?
Your browser tells us the GPU name; our dataset gives its memory and bandwidth with a source for every figure. The engine adds the weights, the context cache and the runtime's overhead, and estimates speed from bandwidth, calibrated on published llama.cpp measurements.
How do I run a model once I've picked it?
Open its page: it shows the best quantization for your machine and the command for llama.cpp or Ollama, plus a link to download it in LM Studio.