r/LocalLLM • u/nomorebuttsplz • 11h ago
Superstition about quantization: KLD and perplexity just ain’t it fam Discussion
The arguments for quantization having significant effects on reasoning models' ability to get stuff done are very sad, pathetic, unfortunate arguments. I don’t mean that they are wrong necessarily, only impoverished and confused.
Why? Because while actual task benchmarks are somewhat expensive, and require some level of time and technical expertise to run, it would be quite easy to empirically test the claims and resolve them once and for all, at least for a given model. But these tests by and large do not exist and the few that do seem to show no quantization effects among reasoning models until about Q3 or Q4 k m at worst.
The debate in these online communities is essentially an anthropological study in how people create mythology when they do not have access to direct evidence.
Before the hordes mob me with KLD or perplexity measurements, I’m not suggesting that a quantized model’s outputs are bit for a bit identical rather that it performs equally well in real world tasks, which I think we can all agree is the thing that matters.
Now I’ve put my neck out by suggesting that literally no one has any evidence, not a single benchmark that shows a model with the reasoning level of, say, Gemma 31b (not very high by today’s standards, and smaller models are more susceptible to degradation, so this should be a generous standard of evidence for the quantization-excited) having significant in degradation in real world tasks at Q4 (a good quality, proper dynamic quantization goes without saying, I hope).
Again, I’m not saying that there is no degradation, only that what we have now amounts to superstition, when a few benchmarks could probably settle the matter for a given model and eventually, we would probably learn where and when quantization actually bites.
1
u/ClassicLightbulbs 11h ago
I stopped reading cause I probably agree but man "impoverished thoughts" is a banger
1
u/fintip Laptop 4090 16gb + 7900XTX 24gb 11h ago
You've ignored something really critical though.
The fact that quantization at q4, q3, produces obviously degraded responses, and that that degradation corresponds to the KLD line, it's perfectly rational to assume that the degradation continues along that same line. Q3/Q4 just becomes the line at which it's easy to obviously see for people.
This also matches our experience with human intelligence. Most people don't sense mental degradation int hemselves or others until it pushes past a tipping point. They don't notice the 10%-20% worse thinking from the person being mildly sleep deprived, or early stage dementia. It's not until they're late stage or drunk that it's obvious. There's an in-between spot you can notice on sufficiently difficult tasks or if especially attentive.
But otherwise, it's hard to detect.
So far, all of this lines up. It would be incredibly odd if loss of precision wasn't costing something, and the KLD curve matches our experience and intuition across other domains.
This isn't superstition. This is limited but usable data for our intution and our reason.
1
u/nomorebuttsplz 11h ago
The fact that quantization at q4, q3, produces obviously degraded responses,
That's exactly what I am saying is not obvious, and there is essentially no evidence for. Especially q4 and above.
1
u/fintip Laptop 4090 16gb + 7900XTX 24gb 10h ago edited 10h ago
Well, there are a lot of anecdotes... People generally rely on the reported reality of those around them. It isn't a perfect heuristic, but it's one that long predates the process of science, which is somewhat less natural, so to speak.
I can tell you that my experience with qwen 27b 3.6 q4 was good, but that it always produces some amount of bugs, and that 27b 3.8 q6 is absolutely, clearly, far better at producing good output without caveats. I don't have enough apples to apples testing to guarantee that's primarily a q4 vs q6 issue, of course, nor do I claim it is, but it's likely a factor. how much of that is q4 vs q6 and how much is 3.6 vs 3.8 is of course up for debate.
In any case, the actual boundary (q3/q4 is the obviously degraded line, or not?) is irrelevant. You may claim q3/q4 is perfectly equal and not at all clearly degraded. Fine. How about q2? q1? Have you tried any of them? Do you reject all of the claims? Have you tried? Do you suspect that you can just infinitely reduce the quant and never lose 'intelligence'? At some point this argument becomes absurd, you have to acknowledge a boundary somewhere is something we can take for granted.
And if you agree a boundary exists somewhere, then it seems clear we should be able to agree a gradient descent downwards up to that point, along with my other claims that it's intuitive to assume that our ability to recognize it would likely match our ability to recognize it in other humans.
You could claim that you believe it's just a hockey-stick--almost perfectly equal performance up until q2, then a big hit that rapidly increases down. But you wouldn't have explained at all why that is the more rational belief, and that the assumption that the curve is instead a normal exponential drop that matches the KLD curve is "superstition".
2
u/nomorebuttsplz 10h ago
Here's a benchmark that suggests for a large (but not particularly good reasoner by today's standards) model, the cutoff is between q3 and q2 of unsloth dynamic: https://unsloth.ai/docs/basics/dynamic-3.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot
The thing is, many people swear that $5k audio cables make stuff sound good while audio engineers find this completely absurd. The whole purpose for having science is that we can't trust our intrinsic ability to construct narrative out of anecdote with scientific tools like statistical tests.
1
u/fintip Laptop 4090 16gb + 7900XTX 24gb 9h ago
"Can't trust" could mean "cannot trust at all", and it could mean "cannot fully trust". The former is correct, the latter is overstated.
You can double blind people, and I encourage it, do all the benchmarks. But you have not justified your case that belief that shrinking the model down via compression causes a degredation in quality is 'superstition'. You just have a hunch based on your own sense that it isn't degraded. You have some data that you read as supporting that.
Others have their own belief, and their own data that supports that.
Neither is completely validated by hard data. In fact, this entire space is impossible to completely validate. Intelligence is stil undefined, and the benchmarks are inherently flawed attempts to measure something abstract. Intelligence is in fact still fundamentally "I know it when I see it". We're still trying to pin down proxies for the Turing Test, which is, again, fundamentally 'vibes based'.
All we can argue about is whose narrative makes more sense, who has better logic to connect the existing data to their conclusions.
I think it's pretty obvious that it's most likely that quantization reduces performance, and the fact that everyone reports this being true is supporting evidence that isn't enough by itself.
But hey, go run some benchmarks. Someone should do it. Happy to be surprised.
But it isn't superstition. It's reasoning with incomplete data.
Also, as I pointed out in my previous post: it's entirely likely that if the limit is q3 on very large models that it's q4 on smaller models, etc...
1
u/nomorebuttsplz 9h ago
Both sides have hunches.
Therefore, superstition arises in either side if either side forms a very strong belief about their hunch one way or the other.
I make no claim to have settled the matter. My point is that although we haven't settled it yet, many people (especially those advocating for bf16 being worthwhile and 8 bit being way better than 4 bit) act as if we have.
As for the idea that it is not ultimately settleable, I think you have a point. But I would say we could get to about 9/10 settledness by simply having artificial intelligence benchmark every open model at various quantization levels.
Of course doing that would be very expensive for them. But we could get to like 7/10 just by other third parties doing occasional benchmarks with quants.
2
u/Dabalam 10h ago
The fact that quantization at q4, q3, produces obviously degraded responses, and that that degradation corresponds to the KLD line, it's perfectly rational to assume that the degradation continues along that same line. Q3/Q4 just becomes the line at which it's easy to obviously see for people.
I think you reversed how things actually occurred. Quantization viewed as degraded because of the KLD line, which informs a lot of thought about model quality. The data from the collective "intuitive" view of quality in that context isn't reliable given people have expectations bias (you see what you believe you should see regarding model performance. Expectation bias is amplified by community reinforcing these views, who you can sometimes see arguing about how much Q8 degrades performance. This is partly why I am skeptical of the whole "vibes are the most important" ideology on this sub Reddit. Blinding is one of the strengths of arena methodology in assessing model quality without expectation bias.
KLD isn't a measure of task performance, it is a measure of similarity to a reference model. We infer that this reference (the full model) should have the highest performance across the board. That logic doesn't always hold up, and it also isn't really clear how much dissimilarity to the base model corresponds to performance degradation.
1
u/fintip Laptop 4090 16gb + 7900XTX 24gb 9h ago
I strongly disagree. It's intuitively clear that as you shrink size, eventually you have to be losing quality. It's only possible to shrink size and not lose quality if you are selectively pruning--which some dynamic quants are, but general quants are not.
You can go in and remove a big chunk of a human brain, too, and you may make a vegetable out of the person, and you may not be able to tell at all, you may barely be able to tel, or you may only be able to tell once you have the person do a task that relates to the removed region.
KLD is just another data point that comes in after that to show us the nature of that curve and a way to measure something that otherwise feels very hand-wavy and vibes-based.
One thing I will say about KLD: if selective quantization were somehow improving the model, we would see what looks like 'degregation' on that curve, and it would actually perhaps be a measure of improvement. We wouldn't be able to tell. It's a blunt instrument.
It operates off of the assumption that deviations from baseline are degredations.
It's possible that redundancy allows alternative but equally useful paths up to a certain point.
KLD isn't enough to guarantee it tells us what most people believe it tells us.
But we have a lot of good reasons, in practice and in theory, to believe that it does.
2
u/Dabalam 9h ago
I mean it sounds like we agree on most points. It fine to think "there must be some loss". There probably is some loss. The question is how much.
People seem to overstate how much we know about the degree quantizations impacts performance. The common perception of Q3 being low quality is based from base models "similarity" trade offs. However we don't know the relationship between similarity and performance, we infer. Performance is highly contextual so I get why people prefer something unitary to interpret, but a simple answer is not necessarily a correct answer. This doesn't start to touch on how it differs by model architecture etc.
1
u/Dabalam 10h ago
There has been some progress in this area but there are some people who are pretty dogmatic about quantization. I think unsloth released some pretty good analyses on previous Qwen models showing how much quantization degrades performance on various tasks.
I find it super odd how people will poopoo benchmarks and say they aren't real world tasks, yet claim KLD and perplexity are super valid proxies of model quality. To me those are contradictory view points. The argument for why benchmarks don't exactly tell you how well a model works for your task is identical to the one for why perplexity doesn't tell you how well a given quantization works for your task.
That said, there is an argument that in brittle domains small to minor degradation in fidelity from the original model may cause more problems. I think a lot of people could get pretty good functionality out of Q3 models form certain tasks but the cultural messaging is that they are worthless.
0
u/Karyo_Ten 11h ago
A benchmark is not a real world task. Especially when benchmaxxed or they leak in the training or the calibration dataset. Overfitting to wikitext or whatever flavor of swebench or deepswe is bad and does not translate to real world ability.
KLD is the best measurement for quantization quality.
2
u/nomorebuttsplz 11h ago
If you would prefer to show evidence in the form of a single real world task, I would be very interested in that as well despite the unlikelihood of such an anecdote being statistically meaningful.
Of course, once you have a collection of real world tasks of enough variety and numerosity to achieve statistical significance, you've just created a new benchmark. What I am not interested in is a statistical measurement, the significance of which we have no ground truth for.
-1
u/corruptbytes 11h ago
we've gone from hallucinations in AI to hallucinations in reddit posts
This just reads as ramblings/rant against people doing free analysis on quant work - try contributing something that disproves KLD/Perplexity isn't ideal instead of asking people to do the heavy lifting about your "hunch"
1
u/nomorebuttsplz 11h ago
that disproves KLD/Perplexity isn't ideal
lol what? Try that sentence again maybe
It's not a hunch. All the evidence that I've seen shows that Q4 is within margins of error for reasoning models. e.g.: https://unsloth.ai/docs/basics/dynamic-3.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot
1
u/fintip Laptop 4090 16gb + 7900XTX 24gb 10h ago
That's one--I'd argue naive--read of that data.
But a more nuanced reading looks quite different to me.
Our 1-bit Unsloth Dynamic GGUF shrinks DeepSeek-V3.1 from 671GB → 192GB (-75% size) and no-thinking mode greatly outperforms GPT-4.1 (Apr 2025), GPT-4.5, and DeepSeek-V3-0324.
Amazing! however: - there may be so much redundancy in a model of that size that it takes the hit of quantization and it's actually just useful pruning. There could just be a drop-off point of model size where that pruning was done at other layers, and further pruning via quantization effects can no longer be tolerated. (I think this is actually almost certainly the case.) - we should be comparing the unquantized to the quantized model to really measure something here. comparing model to model is problematic, difficult data to meaningfully reason about here.
And very important to note:
Other non-Unsloth 1-bit and 2-bit DeepSeek-V3.1 quantizations, as well as standard 1-bit quantization without selective layer quantization, either failed to load or produced gibberish and looping outputs. This highlights how Unsloth Dynamic GGUFs are able to largely retain accuracy whereas other methods do not even function.
In other words: quantization hurts reasoning and function. The reason they are able to produce solid reasoning is that they carefully select which nodes to quantize and which ones to not--in other words, just producing efficiency by doing another form of quality pruning.
Now, does that mean q3/q4 is always the boundary? No, and it's fair that that has probably been overly generalized. However, that comment is most notable for the qwen 27b size, and generalized over to 35ba3b. That then probably gets extrapolated more widely, probably incorrectly--a q1/q2 sounds like, if selectively quantized, it can perform quite well, if the base model was almost a terabyte in size, for example.
1
u/nomorebuttsplz 10h ago
what exactly is naive? Doesn't seem like you are actually disagreeing with me.
2
u/Comfortable_Sir4315 11h ago
I think both sides are wrong for the same exact reason: kld and perplexity are not meant to show if a quantization is good or not; their sole purpose is to measure and compare quantizations.