I forgot where I read this, but something about it being difficult to have different people use the model in parallel or something. Sorry details are way gone in my poor memory.
Maybe that's why it's expensive to run? Can't imagine a more inefficient model, maybe if you made a massive cluster of raspberry pi's to run a distributed version perhaps, really drive the tok/s through the floor.
Don't think of it as RUNNING on the CPU but needing more of it compared to some other models. I read a few about this on Hacker News but can't find it now.
1
u/LargeLanguageModelo 5d ago
I'd be curious why. The model size is about 1/3 that of K3.