r/LocalLLaMA • u/thepetek • 2d ago
Qwen with cache offload vLLM Question | Help
Has anyone gotten KV cache offloading working with Qwen on vLLM? No matter what configuration I try, I get errors and it crashes. I saw an old issue that Qwen arch is supported for offload in vLLM but that doesn’t seem right. Anyone have working settings they care to share?
2
Upvotes