Skip to content

Nemotron 3 Ultra 550B A55B

MoE

NVIDIANVIDIA · 560.5B (55B active) · Mixture of experts

Nemotron 3 Ultra 550B A55B is a 560.5B model from NVIDIA with a 256K-token context. At Q4_K_M with an 8K context it needs about 321.8 GB; 404 GB leaves room for longer chats.

Hugging Face GGUF

Mixture of experts

Parameters: 560.5B · Active: 55B

All 560.5B parameters load into memory, but only 55B work on each token, so it runs at the speed of a much smaller model.

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K0 GB321.8 GB
8K0.1 GB321.8 GB
32K0.4 GB322.1 GB
128K1.5 GB323.2 GB
256K3 GB324.7 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Least hardware that runs it well

Desktop devices with the least memory that give Nemotron 3 Ultra 550B A55B grade A or S at Q4_K_M.

No desktop device in our catalog runs it at grade A or better.

Can I run Nemotron 3 Ultra 550B A55B locally?

Can I run Nemotron 3 Ultra 550B A55B locally?

Yes, if your GPU or Mac has about 321.8 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does Nemotron 3 Ultra 550B A55B need?

About 321.8 GB at Q4_K_M with an 8K context and 544.5 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.