Best local AI models for 8 GB of VRAM
15 open models load in 8 GB of graphics memory at Q4_K_M with an 8K context; 2 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.
Top picks for 8 GB
#1 · ECI 139.5
Qwen3.5 9B
- Memory
- 5.8 GB
- Max context
- 32K
- Speed
- ~50 t/s
#2 · ECI 116.6
Llama 3.1 8B
- Memory
- 5.9 GB
- Max context
- 16K
- Speed
- ~54 t/s
Speeds on RTX 5060 Ti 8GB, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.
Every scored model that fits 8 GB
| # | Model | ECI | Memory | Max context |
|---|---|---|---|---|
| 1 | 139.5 | 5.8 GB | 32K | |
| 2 | 116.6 | 5.9 GB | 16K |
13 more that fit but have no independent score yet
Largest first, with the longest context each can take.
Ministral 3 8B8.9B · 8K
Granite 4.2 8B8.8B · 8K
Gemma 4 E4B8B · 128K
Olmo 3 7B7.3B · 8K
Gemma 4 E2B5.1B · 128K
Qwen3.5 4B4.7B · 128K
Ministral 3 3B3.9B · 32K
Granite 4.2 3B3.7B · 32K
Llama 3.2 3B3.2B · 32K
SmolLM3 3B3.1B · 64K
LFM2.5 2.6B2.7B · 128K
Qwen3.5 2B2.3B · 256K
Qwen3.5 0.8B0.9B · 256K
Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.
What 12 GB adds
1 more scored models fit in 12 GB, the most intelligent of them below.
See the 12 GB guide