Smaller alternatives
Models that need less memory than Kimi K2.7 Code and score about as well or better.
Moonshot AI · 1T (32B active) · Mixture of experts
Kimi K2.7 Code is a 1T model from Moonshot AI with a 256K-token context. At Q4_K_M with an 8K context it needs about 578.8 GB; 725 GB leaves room for longer chats.
Mixture of experts
Parameters: 1T · Active: 32B
All 1T parameters load into memory, but only 32B work on each token, so it runs at the speed of a much smaller model.
| Quant | Bits | Memory | Quality | On your hardware |
|---|---|---|---|---|
| Q2_K | 3.16 | — | Noticeably worse | … |
| Q3_K_M | 4 | — | Some loss | … |
| Q4_K_M | 4.89 | — | Good | … |
| Q5_K_M | 5.7 | — | Good | … |
| Q6_K | 6.56 | — | Near original | … |
| Q8_0 | 8.5 | — | Near original | … |
| F16 | 16 | — | Original | … |
At Q4_K_M. The context cache grows with every token the model keeps in mind.
| Context | Context cache | Total |
|---|---|---|
| 4K | 0.3 GB | 578.5 GB |
| 8K | 0.5 GB | 578.8 GB |
| 32K | 2.1 GB | 580.4 GB |
| 128K | 8.6 GB | 586.8 GB |
| 256K | 17.2 GB | 595.4 GB |
Models that need less memory than Kimi K2.7 Code and score about as well or better.
Scored models of a similar size, side by side on your device.
Desktop devices with the least memory that give Kimi K2.7 Code grade A or S at Q4_K_M.
No desktop device in our catalog runs it at grade A or better.
Yes, if your GPU or Mac has about 578.8 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.
About 578.8 GB at Q4_K_M with an 8K context and 1,017.1 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.