Skip to content

Best local AI models for 32 GB of VRAM

37 open models load in 32 GB of graphics memory at Q4_K_M with an 8K context; 15 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.

Top picks for 32 GB

Speeds on RTX 5090, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.

Every scored model that fits 32 GB

#ModelECIMemoryMax context
1QwenQwen3.8 27B149.417 GB128K
2QwenQwen3.6 27B146.516.5 GB128K
3QwenQwen3.6 35B A3B143.921.2 GB256K
4GoogleGemma 4 31B142.718.8 GB128K
5QwenQwen3.5 35B A3B142.521 GB256K
6GoogleGemma 4 26B A4B141.916.5 GB256K
7QwenQwen3 30B A3B Thinking 2507139.618.3 GB128K
8QwenQwen3.5 9B139.55.8 GB256K
9OpenAIgpt-oss 20b137.811.3 GB128K
10QwenQwen3 30B A3B 2507137.418.3 GB128K
11Mistral AIMagistral Small 2506133.214.9 GB32K
12Mistral AIMistral Small 3.2 24B131.714.9 GB64K
13Mistral AIMagistral Small 2509131.414.9 GB64K
14MicrosoftPhi-4 14B130.410.1 GB16K
15MetaLlama 3.1 8B116.65.9 GB128K

22 more that fit but have no independent score yet

Largest first, with the longest context each can take.

Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.

What 48 GB adds

1 more scored models fit in 48 GB, the most intelligent of them below.

See the 48 GB guide

Devices with 32 GB