Best local AI models for 48 GB of VRAM
39 open models load in 48 GB of graphics memory at Q4_K_M with an 8K context; 16 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.
Top picks for 48 GB
#1 · ECI 149.4
Qwen3.8 27B
- Memory
- 17 GB
- Max context
- 256K
- Speed
- ~28 t/s
#2 · ECI 146.5
Qwen3.6 27B
- Memory
- 16.5 GB
- Max context
- 256K
- Speed
- ~29 t/s
#3 · ECI 143.9
Qwen3.6 35B A3B
- Memory
- 21.2 GB
- Max context
- 256K
- Speed
- ~259 t/s
Speeds on RTX A6000, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.
Every scored model that fits 48 GB
| # | Model | ECI | Memory | Max context |
|---|---|---|---|---|
| 1 | 149.4 | 17 GB | 256K | |
| 2 | 146.5 | 16.5 GB | 256K | |
| 3 | 143.9 | 21.2 GB | 256K | |
| 4 | 142.7 | 18.8 GB | 256K | |
| 5 | 142.5 | 21 GB | 256K | |
| 6 | 141.9 | 16.5 GB | 256K | |
| 7 | 139.6 | 18.3 GB | 256K | |
| 8 | 139.5 | 5.8 GB | 256K | |
| 9 | 137.8 | 11.3 GB | 128K | |
| 10 | 137.4 | 18.3 GB | 256K | |
| 11 | 133.2 | 14.9 GB | 32K | |
| 12 | 131.7 | 14.9 GB | 128K | |
| 13 | 131.4 | 14.9 GB | 128K | |
| 14 | 130.4 | 10.1 GB | 16K | |
| 15 | 127.3 | 42.4 GB | 16K | |
| 16 | 116.6 | 5.9 GB | 128K |
23 more that fit but have no independent score yet
Largest first, with the longest context each can take.
Qwen3 Coder Next79.7B · 64K
Kimi Linear 48B A3B49.1B · 1M
Olmo 3.1 32B Think32.2B · 64K
Nemotron 3.5 Lightning 30B A3B31.6B · 256K
GLM 4.7 Flash31.2B · 128K
North Mini Code 1.030.5B · 256K
Granite 4.2 30B29.3B · 64K
Devstral Small 2 24B24B · 128K
Ministral 3 14B14B · 128K
Gemma 4 12B12B · 256K
Ministral 3 8B8.9B · 256K
Granite 4.2 8B8.8B · 128K
Gemma 4 E4B8B · 128K
Olmo 3 7B7.3B · 64K
Gemma 4 E2B5.1B · 128K
Qwen3.5 4B4.7B · 256K
Ministral 3 3B3.9B · 256K
Granite 4.2 3B3.7B · 128K
Llama 3.2 3B3.2B · 128K
SmolLM3 3B3.1B · 64K
LFM2.5 2.6B2.7B · 128K
Qwen3.5 2B2.3B · 256K
Qwen3.5 0.8B0.9B · 256K
Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.