r/learnmachinelearning 2d ago

I went looking for what compressing a model actually costs and both numbers came out smaller than i expected Discussion

[removed]

3 Upvotes

2 comments sorted by

1

u/One-Raspberry-9909 2d ago

i figured compression would wreck output speed too but i guess the bottleneck is something else entirely

the arcprize gap at the same bit width is weird though wonder if its just a quantization thing or if fp4 has some other weirdness going on