Skip to content

LFM2.5 2.6B

Liquid AILiquid AI · 2.7B · Dense

LFM2.5 2.6B is a 2.7B model from Liquid AI with a 128K-token context. At Q4_K_M with an 8K context it needs about 2 GB; 4 GB leaves room for longer chats.

Hugging Face GGUF

Quantization options

QuantMemoryOn your hardware
Q2_K—…
Q3_K_M—…
Q4_K_M—…
Q5_K_M—…
Q6_K—…
Q8_0—…
F16—…

Memory by context length

At Q4_K_M. The context cache grows with every token the model keeps in mind.

ContextContext cacheTotal
4K0.1 GB1.9 GB
8K0.1 GB2 GB
32K0.5 GB2.4 GB
128K2 GB3.9 GB

Share a measured speed

Run one of these and paste the whole output below (or just the tokens per second):

llama-bench -m model.gguf
ollama run model --verbose

Published anonymously and kept. We store no account or address, only a daily-changing hash for a limit of ten submissions a day.

Next steps

Smaller alternatives

Models that need less memory than LFM2.5 2.6B and score about as well or better.

No scored model is both smaller and about as intelligent as LFM2.5 2.6B.

Can I run LFM2.5 2.6B locally?

Can I run LFM2.5 2.6B locally?

Yes, if your GPU or Mac has about 2 GB free for it at Q4_K_M. Open this page on that computer to see the grade for your exact hardware, then start it with llama.cpp, Ollama or LM Studio.

How much memory does LFM2.5 2.6B need?

About 2 GB at Q4_K_M with an 8K context and 3.1 GB at Q8_0. Lower quantizations fit smaller cards with some loss in quality; longer contexts add to the total.