Best local AI models for 24 GB of VRAM
35 open models load in 24 GB of graphics memory at Q4_K_M with an 8K context; 15 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.
Top picks for 24 GB
#1 · ECI 149.4
Qwen3.8 27B
- Memory
- 17 GB
- Max context
- 64K
- Speed
- ~37 t/s
#2 · ECI 146.5
Qwen3.6 27B
- Memory
- 16.5 GB
- Max context
- 64K
- Speed
- ~39 t/s
#3 · ECI 143.9
Qwen3.6 35B A3B
- Memory
- 21.2 GB
- Max context
- 64K
- Speed
- ~340 t/s
Speeds on RTX 3090 Ti, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.
Every scored model that fits 24 GB
| # | Model | ECI | Memory | Max context |
|---|---|---|---|---|
| 1 | 149.4 | 17 GB | 64K | |
| 2 | 146.5 | 16.5 GB | 64K | |
| 3 | 143.9 | 21.2 GB | 64K | |
| 4 | 142.7 | 18.8 GB | 32K | |
| 5 | 142.5 | 21 GB | 64K | |
| 6 | 141.9 | 16.5 GB | 256K | |
| 7 | 139.6 | 18.3 GB | 32K | |
| 8 | 139.5 | 5.8 GB | 256K | |
| 9 | 137.8 | 11.3 GB | 128K | |
| 10 | 137.4 | 18.3 GB | 32K | |
| 11 | 133.2 | 14.9 GB | 32K | |
| 12 | 131.7 | 14.9 GB | 32K | |
| 13 | 131.4 | 14.9 GB | 32K | |
| 14 | 130.4 | 10.1 GB | 16K | |
| 15 | 116.6 | 5.9 GB | 128K |
20 more that fit but have no independent score yet
Largest first, with the longest context each can take.
Olmo 3.1 32B Think32.2B · 32K
GLM 4.7 Flash31.2B · 64K
North Mini Code 1.030.5B · 128K
Granite 4.2 30B29.3B · 16K
Devstral Small 2 24B24B · 32K
Ministral 3 14B14B · 64K
Gemma 4 12B12B · 256K
Ministral 3 8B8.9B · 128K
Granite 4.2 8B8.8B · 64K
Gemma 4 E4B8B · 128K
Olmo 3 7B7.3B · 64K
Gemma 4 E2B5.1B · 128K
Qwen3.5 4B4.7B · 256K
Ministral 3 3B3.9B · 128K
Granite 4.2 3B3.7B · 128K
Llama 3.2 3B3.2B · 128K
SmolLM3 3B3.1B · 64K
LFM2.5 2.6B2.7B · 128K
Qwen3.5 2B2.3B · 256K
Qwen3.5 0.8B0.9B · 256K
Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.