Skip to content
All devices

48 GB VRAM · 864 GB/s · CUDA

AI models for L40S

Here are the best open models for each job on the L40S, and every model's grade below.

Your hardware

…

Memory
GB
Bandwidth
GB/s
Context

 

Most intelligent on your …

Intelligence (ECI) of the most intelligent models that load · colour shows how well they run

Expected speed

Tokens per second on your … for the most intelligent models that run well

Models available with more memory

Open models that run well (grade B or better) at each memory size

Scores from Oct 8, 2026 · 0 new models in the last two months

Every model on the L40S

Q4_K_M, 8K context, 48 GB of memory. Change your setup above for exact numbers.

GradeModelMemorySpeed
SmoonshotKimi Linear 48B A3B 49.1B28.4 GB~303 tok/s
SqwenQwen3.6 35B A3B 36B21.2 GB~291 tok/s
SqwenQwen3.5 35B A3B 36B21 GB~295 tok/s
SnvidiaNemotron 3.5 Lightning 30B A3B 31.6B24.1 GB~231 tok/s
SzaiGLM 4.7 Flash 31.2B17.8 GB~286 tok/s
SqwenQwen3 30B A3B Thinking 2507 30.5B18.3 GB~235 tok/s
SqwenQwen3 30B A3B Instruct 2507 30.5B18.3 GB~235 tok/s
ScohereNorth Mini Code 1.0 30.5B18.2 GB~251 tok/s
SgoogleGemma 4 26B A4B 25.8B16.5 GB~202 tok/s
Sopenaigpt-oss 20b 20.9B11.3 GB~269 tok/s
SmicrosoftPhi-4 14B 14.7B10.1 GB~58 tok/s
SmistralMinistral 3 14B 14B9.2 GB~64 tok/s
SgoogleGemma 4 12B 12B7.4 GB~75 tok/s
SqwenQwen3.5 9B 9.7B5.8 GB~97 tok/s
SmistralMinistral 3 8B 8.9B6.2 GB~98 tok/s
SibmGranite 4.2 8B 8.8B6.7 GB~91 tok/s
SmetaLlama 3.1 8B Instruct 8B5.9 GB~104 tok/s
SgoogleGemma 4 E4B 8B5.1 GB~112 tok/s
Sai2Olmo 3 7B Instruct 7.3B7 GB~86 tok/s
SgoogleGemma 4 E2B 5.1B3.3 GB~180 tok/s
SqwenQwen3.5 4B 4.7B3.1 GB~197 tok/s
SmistralMinistral 3 3B 3.9B3.1 GB~219 tok/s
SibmGranite 4.2 3B 3.7B3.1 GB~213 tok/s
SmetaLlama 3.2 3B Instruct 3.2B3.1 GB~227 tok/s
ShuggingfaceSmolLM3 3B 3.1B2.6 GB~255 tok/s
SliquidLFM2.5 2.6B 2.7B2 GB~323 tok/s
SqwenQwen3.5 2B 2.3B1.6 GB~425 tok/s
SqwenQwen3.5 0.8B 0.9B0.9 GB~971 tok/s
Aai2Olmo 3.1 32B Think 32.2B19.7 GB~28 tok/s
AgoogleGemma 4 31B 31.3B18.8 GB~29 tok/s
AibmGranite 4.2 30B 29.3B19.1 GB~30 tok/s
AqwenQwen3.8 27B 27.8B17 GB~32 tok/s
AqwenQwen3.6 27B 27.8B16.5 GB~33 tok/s
AmistralDevstral Small 2 24B 24B14.9 GB~38 tok/s
AmistralMagistral Small 2509 24B14.9 GB~38 tok/s
AmistralMistral Small 3.2 24B 24B14.9 GB~38 tok/s
AmistralMagistral Small 2506 23.6B14.9 GB~38 tok/s
BqwenQwen3 Coder Next 79.7B45.7 GB~294 tok/s
BmetaLlama 3.3 70B Instruct 70.6B42.4 GB~13 tok/s
CmistralMistral Small 4 119B 119.4B68.1 GB~33 tok/s
Copenaigpt-oss 120b 116.8B59 GB~63 tok/s
CmetaLlama 4 Scout 17B 16E 108.6B62.7 GB~14 tok/s
32 models do not fit on the L40S
GradeModelMemorySpeed
FmoonshotKimi K3 2.8T1,584.5 GB—
FdeepseekDeepSeek V4 Pro 0813 1.7T886.8 GB—
FdeepseekDeepSeek V4 Pro 1.6T886.8 GB—
FmoonshotKimi K2.6 1T578.8 GB—
FmoonshotKimi K2.7 Code 1T578.8 GB—
FmoonshotKimi K2.5 1T579.4 GB—
FmoonshotKimi K2 Thinking 1T579.4 GB—
FmoonshotKimi K2 Instruct 1T579 GB—
FtInkling 952.4B534 GB—
FdeepseekDeepSeek V4.1 Flash 763.2B435.8 GB—
FzaiGLM 5.1 753.9B429.2 GB—
FzaiGLM 5 753.9B425.7 GB—
FzaiGLM 5.2 753.3B424.5 GB—
FdeepseekDeepSeek V3.2 Exp 685.4B391.5 GB—
FdeepseekDeepSeek V3.2 685.4B378.5 GB—
FdeepseekDeepSeek V3.1 684.5B378.4 GB—
FnvidiaNemotron 3 Ultra 550B A55B 560.5B321.8 GB—
FminimaxMiniMax M3 427B244.6 GB—
FqwenQwen3.5 397B A17B 403.4B227.9 GB—
FzaiGLM 4.7 358.3B204.8 GB—
FzaiGLM 5.3 Flash 321.3B176.1 GB—
FdeepseekDeepSeek V4 Flash 0731 304.2B161 GB—
FdeepseekDeepSeek V4 Flash 290.9B166.7 GB—
FtInkling Small 266B149 GB—
FqwenQwen3 235B A22B Thinking 2507 235.1B134.2 GB—
FqwenQwen3 235B A22B Instruct 2507 235.1B134.2 GB—
FminimaxMiniMax M2.5 228.7B131.1 GB—
FminimaxMiniMax M2.7 228.7B131.5 GB—
FcohereCommand A+ 218.8B126.7 GB—
FqwenQwen3.8 Flash Next 180B111.9 GB—
FmistralMistral Medium 3.5 128B 127.7B72.8 GB—
FqwenQwen3.5 122B A10B 125.1B71.8 GB—

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps