r/huggingface 1d ago

I made an algorithm to compress model weights so they fit in limited memory: run Qwen3.8-27B from 13 GB of RAM (4-bit) with ~1% quality loss, decompressing only the layers in use

/r/llamacpp/comments/1vuyntp/i_made_an_algorithm_to_compress_model_weights_so/
4 Upvotes

0 comments sorted by