Skip to content

DeepSeek V4 Flash 0731

MoE

DeepSeekDeepSeek · 304.2B (13B active) · Mixture of experts

DeepSeek V4 Flash 0731 is a 304.2B model from DeepSeek with a 1M-token context. At Q4_K_M with an 8K context it needs about 161 GB; 203 GB leaves room for longer chats.

Hugging Face GGUF

Mixture of experts

Parameters: 304.2B · Active: 13B

All 304.2B parameters load into memory, but only 13B work on each token, so it runs at the speed of a much smaller model.

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K0.3 GB160.7 GB
8K0.7 GB161 GB
32K2.7 GB163 GB
128K10.8 GB171.1 GB
256K21.5 GB181.8 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Smaller alternatives

Models that need less memory than DeepSeek V4 Flash 0731 and score about as well or better.

No scored model is both smaller and about as intelligent as DeepSeek V4 Flash 0731.

Least hardware that runs it well

Desktop devices with the least memory that give DeepSeek V4 Flash 0731 grade A or S at Q4_K_M.

Can I run DeepSeek V4 Flash 0731 locally?

Can I run DeepSeek V4 Flash 0731 locally?

Yes, if your GPU or Mac has about 161 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does DeepSeek V4 Flash 0731 need?

About 161 GB at Q4_K_M with an 8K context and 282.4 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.