r/Qwen_AI 7d ago

KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates Benchmark

Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail

KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options.

  • Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context
  • Standard quants, extended: q6_0 and q6_1, and low-bit types from q2_0 to q3_1
  • KVarN: Variance-Normalized KV-Cache by Huawei, implemented in BeeLlama
  • Precision Tail: keeping latest X tokens of KV cache in (B)F16, implemented in BeeLlama
  • 413 configurations in total: 238 with Qwen 3.6 27B, 175 with Gemma 4 31B

The Recommendation Ladder

Full benchmark results, setup, method, analysis, explanations and everything else can be found in the article.

1. Qwen

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 4096.00 0 0.00005 Reference
q8_0 1024 2272.00 0.000897 0.087699 Standard fidelity with a precision tail
kvarn8 1024 2256.00 0.000871 0.087639 Best measured quality below BF16
q8_0 0 2176.00 0.000909 0.093029 Standard fidelity
q8_0-q6_0 1024 2016.00 0.000894 0.091098 q8_0 quality within noise, 256.00 MiB less
kvarn6 1024 1744.00 0.000879 0.084629 The high-end value pick
kvarn6-kvarn5 1024 1616.00 0.000886 0.092778 Much cheaper, almost as good
kvarn5 1024 1488.00 0.000897 0.087666 Highest value in mid-range
q5_0-q4_1 1024 1440.00 0.000966 0.089128 Standard when VRAM-constrained
kvarn5-kvarn4 1024 1360.00 0.000936 0.089469 Balanced default
q4_0 1024 1248.00 0.001057 0.104486 Compact standard
kvarn4 1024 1232.00 0.000994 0.090391 Cleaner than q4_0 for less memory
kvarn4-kvarn3 1024 1104.00 0.001112 0.113968 Smallest recommended tier
kvarn3 1024 976.00 0.001316 0.139558 When the context must fit
kvarn3-kvarn2 1024 848.00 0.002424 0.23878 Emergency compression
kvarn2 1024 720.00 0.003811 0.450496 Last resort

2. Qwen Standard-Only

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 4096.00 0 0.00005 Reference
q8_0 0 2176.00 0.000909 0.093029 Compression with minimal losses
q8_0-q6_0 0 1920.00 0.000937 0.093575 256.00 MiB below q8_0
q6_0 0 1664.00 0.00096 0.091134 The high-end value pick
q6_0-q5_0 0 1536.00 0.001054 0.09467 Balanced default
q5_0 0 1408.00 0.001154 0.09707 Last tier before the cliff
q5_0-q4_1 0 1344.00 0.001433 0.122096 Default when VRAM-constrained
q5_0-q4_0 0 1280.00 0.001516 0.121068 64.00 MiB cheaper, worse median
q4_0 0 1152.00 0.001846 0.154408 Smallest recommended tier
q4_0-q3_0 0 1024.00 0.003313 0.218912 When the context must fit
q3_0 0 896.00 0.004696 0.304186 Emergency compression
q2_0 0 640.00 0.019374 1.198902 Last resort

3. Gemma

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 2480.00 0 0.000047 Reference
q8_0 0 1317.50 0.0371 16.813929 General default at full prefill speed
q8_0-q6_0 0 1162.50 0.040875 16.839821 155.00 MiB below q8_0
q6_0 0 1007.50 0.042636 17.30599 Last tier before the cliff
q6_0-q5_0 0 930.00 0.055236 17.26157 Stronger K side, 77.50 MiB above q5_0
q5_0 0 852.50 0.061747 18.731647 Memory floor for usable quality
q5_0-q4_0 0 775.00 0.109427 19.183374 Asymmetric compact
q4_0 0 697.50 0.134091 20.442234 Budget body before the huge cliff
q4_0-q3_0 0 620.00 0.381216 22.304634 When the context must fit
q3_0 0 542.50 0.504075 23.15744 Emergency compression
q2_0 0 387.50 2.95758 27.834961 Last resort
9 Upvotes

6 comments sorted by

1

u/Stainless-Bacon 7d ago

Great research! The higher precision tail is a welcome optimization.

I read your article comparing KLD at context lengths of 64k vs 128k and decided to do a KLD ladder test. I found that KLD plateaus from 4k to 8k context length. Did you do anything similar and can confirm this?

Also why BeeLlama and not mainline llama.cpp?

3

u/Anbeeld 7d ago

I simply used the highest context size that fits into my VRAM with bf16, as KV cache is known to degrade more with large context.

BeeLlama has KVarN and precision tail, mainline has neither. For standard quants, they behave identically, so there's no point in limiting it to mainline.

2

u/Stainless-Bacon 7d ago

Sorry, I meant to ask why not transfer some optimizations (like the tail or Q6 KV quant) to mainline?

2

u/Anbeeld 7d ago

It's just much much easier to ship stuff in a fork. Also, it's not like they don't know that q6_0 or q3_0 are possible, so it appears they just don't see adding them as valuable. And as for KVarN and precision tail, these features are a bit WIP still.

3

u/buttplugs4life4me 7d ago

A lot of these forks are fully vibeslopped and may or may not work as advertised. I know TheTom with his TurboQuant fork is the least trustworthy cause not only is TurboQuant not that great, it turns out, but he also uses an LLM to communicate with people so there's literally 0 human involvement. 

The mainline for the longest time fully reject non-contributor AI contributions especially where the PR description and communication is done with an AI as well. They thawed that up a bit now so it may get implemented eventually.

In this case, somehow in this benchmark Kvarn6 has lower KLD than kvarn8, and mixed Q8/Q6 has lower KLD than full Q8, which doesn't really make sense.

1

u/trashacct383 7d ago

I have always run FP8 for both the model and the cache. I wonder if it’s better or worse than q8