Skip to content

Best local AI models for 16 GB of VRAM

23 open models load in 16 GB of graphics memory at Q4_K_M with an 8K context; 7 of them have an independent intelligence score. Here they are ranked, with the longest context each can take.

Top picks for 16 GB

Speeds on RTX 5080, the newest NVIDIA card with this much memory. Max context is the longest that still loads without spilling into system RAM.

Every scored model that fits 16 GB

#ModelECIMemoryMax context
1QwenQwen3.5 9B139.55.8 GB256K
2OpenAIgpt-oss 20b137.811.3 GB128K
3Mistral AIMagistral Small 2506133.214.9 GB8K
4Mistral AIMistral Small 3.2 24B131.714.9 GB8K
5Mistral AIMagistral Small 2509131.414.9 GB8K
6MicrosoftPhi-4 14B130.410.1 GB16K
7MetaLlama 3.1 8B116.65.9 GB64K

16 more that fit but have no independent score yet

Largest first, with the longest context each can take.

Memory at Q4_K_M with an 8K context, including the runtime’s own buffers. A smaller quant fits more; a longer context needs more.

What 24 GB adds

8 more scored models fit in 24 GB, the most intelligent of them below.

See the 24 GB guide

Devices with 16 GB