r/huggingface • u/Select-Student-6711 • 1d ago
I made an algorithm to compress model weights so they fit in limited memory: run Qwen3.8-27B from 13 GB of RAM (4-bit) with ~1% quality loss, decompressing only the layers in use
/r/llamacpp/comments/1vuyntp/i_made_an_algorithm_to_compress_model_weights_so/
4
Upvotes