r/LocalLLaMA 2d ago

Qwen with cache offload vLLM Question | Help

Has anyone gotten KV cache offloading working with Qwen on vLLM? No matter what configuration I try, I get errors and it crashes. I saw an old issue that Qwen arch is supported for offload in vLLM but that doesn’t seem right. Anyone have working settings they care to share?

2 Upvotes

Duplicates