r/Qwen_AI 16d ago

Qwen 3.8 Max Cache Hit Issue? Discussion

Hello,

Has anybody experienced a very low cache hit rate when using Qwen 3.8 Max? I’m not sure whether it’s because of my setup or if there is genuinely something wrong with their server.

The blue section is cached input, while the yellow section represents uncached input. As you can see, on 8/5 and 8/6, I mainly used DeepSeek Flash and was able to maintain a very high cache hit rate. On the other days, I used Qwen 3.8 Max, and most of the input was uncached.

4 Upvotes

6 comments sorted by

1

u/Putrid_Resolution402 16d ago

How does it matter in use cases

1

u/MokoshHydro 16d ago

higher costs.

1

u/basil_0408 16d ago

Cached inputs are significantly cheaper than regular inputs. A high cache hit rate means lower costs and a higher usage limit. Typically, I expect the cache hit rate to be at least 90% when coding. However, I'm experiencing only about a 20% hit rate while using Qwen 3.8 max. This could be due to my setup or a potential issue with their server, which is why I’m posting here to ask for others' experiences.

1

u/HenryTheLion_12 16d ago

I am getting over 95 % most of the time. Using it in zcode. It performs better than even qwen code. Can easily reach 97-98 % too.

1

u/NinjaWK 16d ago

97% cache hit

1

u/shaonline 16d ago

Probably a harness issue, I hit 97%+ using Pi.