Skip to content

Qwen3 30B A3B Instruct 2507

MoE

QwenQwen · 30.5B (3.3B active) · Mixture of experts

Qwen3 30B A3B Instruct 2507 is a 30.5B model from Qwen with a 256K-token context. At Q4_K_M with an 8K context it needs about 18.3 GB; 24 GB leaves room for longer chats.

Hugging Face GGUF

Mixture of experts

Parameters: 30.5B · Active: 3.3B

All 30.5B parameters load into memory, but only 3.3B work on each token, so it runs at the speed of a much smaller model.

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K0.4 GB18 GB
8K0.8 GB18.3 GB
32K3 GB20.6 GB
128K12 GB29.6 GB
256K24 GB41.6 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Can I run Qwen3 30B A3B Instruct 2507 locally?

Can I run Qwen3 30B A3B Instruct 2507 locally?

Yes, if your GPU or Mac has about 18.3 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does Qwen3 30B A3B Instruct 2507 need?

About 18.3 GB at Q4_K_M with an 8K context and 31.3 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.