Best local AI models for 12 GB of VRAM
18 open models load in 12 GB of graphics memory at Q4_K_M with an 8K context; 3 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.
Top picks for 12 GB
#1 · ECI 139.5
Qwen3.5 9B
- Memory
- 5.8 GB
- Max context
- 128K
- Speed
- ~76 t/s
#2 · ECI 130.4
Phi-4 14B
- Memory
- 10.1 GB
- Max context
- 8K
- Speed
- ~45 t/s
#3 · ECI 116.6
Llama 3.1 8B
- Memory
- 5.9 GB
- Max context
- 32K
- Speed
- ~81 t/s
Speeds on RTX 5070, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.
Every scored model that fits 12 GB
| # | Model | ECI | Memory | Max context |
|---|---|---|---|---|
| 1 | 139.5 | 5.8 GB | 128K | |
| 2 | 130.4 | 10.1 GB | 8K | |
| 3 | 116.6 | 5.9 GB | 32K |
15 more that fit but have no independent score yet
Largest first, with the longest context each can take.
Ministral 3 14B14B · 16K
Gemma 4 12B12B · 128K
Ministral 3 8B8.9B · 32K
Granite 4.2 8B8.8B · 32K
Gemma 4 E4B8B · 128K
Olmo 3 7B7.3B · 32K
Gemma 4 E2B5.1B · 128K
Qwen3.5 4B4.7B · 256K
Ministral 3 3B3.9B · 64K
Granite 4.2 3B3.7B · 64K
Llama 3.2 3B3.2B · 64K
SmolLM3 3B3.1B · 64K
LFM2.5 2.6B2.7B · 128K
Qwen3.5 2B2.3B · 256K
Qwen3.5 0.8B0.9B · 256K
Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.
What 16 GB adds
4 more scored models fit in 16 GB, the most intelligent of them below.
See the 16 GB guide