Skip to content

Mistral Medium 3.5 128B

Vision

Mistral AIMistral AI · 127.7B · Dense

Mistral Medium 3.5 128B is a 127.7B model from Mistral AI with a 256K-token context. At Q4_K_M with an 8K context it needs about 72.8 GB; 93 GB leaves room for longer chats.

Hugging Face GGUF

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K1.4 GB71.4 GB
8K2.8 GB72.8 GB
32K11 GB81.1 GB
128K44 GB114.1 GB
256K88 GB158.1 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Least hardware that runs it well

Desktop devices with the least memory that give Mistral Medium 3.5 128B grade A or S at Q4_K_M.

No desktop device in our catalog runs it at grade A or better.

Can I run Mistral Medium 3.5 128B locally?

Can I run Mistral Medium 3.5 128B locally?

Yes, if your GPU or Mac has about 72.8 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does Mistral Medium 3.5 128B need?

About 72.8 GB at Q4_K_M with an 8K context and 126.8 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.