Best local AI models for 16 GB of VRAM
23 open models load in 16 GB of graphics memory at Q4_K_M with an 8K context; 7 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.
Top picks for 16 GB
#1 · ECI 139.5
Qwen3.5 9B
- Memory
- 5.8 GB
- Max context
- 256K
- Speed
- ~108 t/s
#2 · ECI 137.8
gpt-oss 20b
- Memory
- 11.3 GB
- Max context
- 128K
- Speed
- ~299 t/s
#3 · ECI 133.2
Magistral Small 2506
- Memory
- 14.9 GB
- Max context
- 8K
- Speed
- ~42 t/s
Speeds on RTX 5080, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.
Every scored model that fits 16 GB
| # | Model | ECI | Memory | Max context |
|---|---|---|---|---|
| 1 | 139.5 | 5.8 GB | 256K | |
| 2 | 137.8 | 11.3 GB | 128K | |
| 3 | 133.2 | 14.9 GB | 8K | |
| 4 | 131.7 | 14.9 GB | 8K | |
| 5 | 131.4 | 14.9 GB | 8K | |
| 6 | 130.4 | 10.1 GB | 16K | |
| 7 | 116.6 | 5.9 GB | 64K |
16 more that fit but have no independent score yet
Largest first, with the longest context each can take.
Devstral Small 2 24B24B · 8K
Ministral 3 14B14B · 32K
Gemma 4 12B12B · 256K
Ministral 3 8B8.9B · 64K
Granite 4.2 8B8.8B · 32K
Gemma 4 E4B8B · 128K
Olmo 3 7B7.3B · 64K
Gemma 4 E2B5.1B · 128K
Qwen3.5 4B4.7B · 256K
Ministral 3 3B3.9B · 64K
Granite 4.2 3B3.7B · 128K
Llama 3.2 3B3.2B · 64K
SmolLM3 3B3.1B · 64K
LFM2.5 2.6B2.7B · 128K
Qwen3.5 2B2.3B · 256K
Qwen3.5 0.8B0.9B · 256K
Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.
What 24 GB adds
8 more scored models fit in 24 GB, the most intelligent of them below.
See the 24 GB guide