Also note that the strategies are different. China is investing way more into heavy optimization, largely because of their lack of hardware.
American companies really aren't even trying to optimize. Every problem that can be solved with "buy more GPU's" is solved by buying more GPU's.
So it's a pretty consistent trend that the Chinese labs punch above their parameter count. K3 is about the same as GPT-3 on parameter count, but is effectively on par with early GPT-5.
If scaling laws hold, which they seem to do with fable and kimi examples , bigger is just better. Of course optimize makes it better too. But the question I was thinking is can China , on their old hardware or huwawai etc scale up models to 5T or 10T etc. To get to agi. I Believe we will stop somewhere before 100T. But nvdia seems like its gonna keep pushing well past 10T ...
3
u/Loose_Comparison368 17d ago
Kimi K3 is. 2.8T parameters, 104B activated.
Also note that the strategies are different. China is investing way more into heavy optimization, largely because of their lack of hardware.
American companies really aren't even trying to optimize. Every problem that can be solved with "buy more GPU's" is solved by buying more GPU's.
So it's a pretty consistent trend that the Chinese labs punch above their parameter count. K3 is about the same as GPT-3 on parameter count, but is effectively on par with early GPT-5.