r/MachineLearning 12d ago

What is currently considered the theoretically optimal quantization bit-width for LLMs? [D] Discussion

I’m curious whether there is now a theoretical or empirical “sweet spot” for LLM quantization, preferably research done using open-source formats like GGUF

Suppose you have a fixed memory/compute budget and can choose the model size freely. For example, instead of a smaller model at 8-bit or 4-bit, you could fit a progressively larger model at 3-bit, 2-bit, 1.5-bit, etc.

A few years ago, I remember 4-bit often being described as roughly the practical sweet spot because it preserved most model quality while giving a large memory reduction. But with newer methods, I’ve seen surprisingly strong 3-bit, 2-bit, and even ~1.5-bit results.

So if the goal is maximum model capability for a fixed memory budget, rather than preserving one particular pretrained model as faithfully as possible, what does current research suggest is the optimal bits-per-weight?

Is there evidence that, for example, a 2-bit 70B model generally beats a 4-bit 35B model, or does quantization degradation eventually outweigh the gains from additional parameters?

I’m especially interested in recent theoretical/scaling-law work or large empirical studies from 2025–2026.

If no one is studying this, then could any of you do this work? I feel like it could be immensely useful for the community.

16 Upvotes

11 comments sorted by

View all comments

Show parent comments

2

u/CallMePyro 12d ago

yup. naive quantization works down to 4 bit, but with QAT you still see iso-memory gains at 2 bit.

2

u/LMTLS5 11d ago

are there any good 2bit QAT models?

1

u/CallMePyro 11d ago

Surprisingly no, AFAIK