Skip to content

Gemma 4 31B

Vision

GoogleGoogle · 31.3B · Dense

Gemma 4 31B is a 31.3B model from Google with a 256K-token context. At Q4_K_M with an 8K context it needs about 18.8 GB; 25 GB leaves room for longer chats.

Hugging Face GGUF

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K1.1 GB18.5 GB
8K1.4 GB18.8 GB
32K3.3 GB20.6 GB
128K10.8 GB28.1 GB
256K20.8 GB38.1 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Smaller alternatives

Models that need less memory than Gemma 4 31B and score about as well or better.

No scored model is both smaller and about as intelligent as Gemma 4 31B.

Can I run Gemma 4 31B locally?

Can I run Gemma 4 31B locally?

Yes, if your GPU or Mac has about 18.8 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does Gemma 4 31B need?

About 18.8 GB at Q4_K_M with an 8K context and 32.1 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.