Smaller alternatives
Models that need less memory than Qwen3 235B A22B Thinking 2507 and score about as well or better.
Qwen · 235.1B (22B active) · Mixture of experts
Qwen3 235B A22B Thinking 2507 is a 235.1B model from Qwen with a 256K-token context. At Q4_K_M with an 8K context it needs about 134.2 GB; 169 GB leaves room for longer chats.
Mixture of experts
Parameters: 235.1B · Active: 22B
All 235.1B parameters load into memory, but only 22B work on each token, so it runs at the speed of a much smaller model.
| Quant | Bits | Memory | Quality | On your hardware |
|---|---|---|---|---|
| Q2_K | 3.16 | — | Noticeably worse | … |
| Q3_K_M | 4 | — | Some loss | … |
| Q4_K_M | 4.89 | — | Good | … |
| Q5_K_M | 5.7 | — | Good | … |
| Q6_K | 6.56 | — | Near original | … |
| Q8_0 | 8.5 | — | Near original | … |
| F16 | 16 | — | Original | … |
At Q4_K_M. The context cache grows with every token the model keeps in mind.
| Context | Context cache | Total |
|---|---|---|
| 4K | 0.7 GB | 133.4 GB |
| 8K | 1.5 GB | 134.2 GB |
| 32K | 5.9 GB | 138.6 GB |
| 128K | 23.5 GB | 156.2 GB |
| 256K | 47 GB | 179.7 GB |
Models that need less memory than Qwen3 235B A22B Thinking 2507 and score about as well or better.
Scored models of a similar size, side by side on your device.
Desktop devices with the least memory that give Qwen3 235B A22B Thinking 2507 grade A or S at Q4_K_M.
Yes, if your GPU or Mac has about 134.2 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.
About 134.2 GB at Q4_K_M with an 8K context and 234.5 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.
Your privacy
We use PostHog to understand how the site is used and improve it. If you accept, it stores an identifier in your browser and records your session (text fields stay hidden). If you reject, we only count the visit anonymously, without cookies. You can change this any time from the footer.