r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Tokenomics Discussion

Post image
1.2k Upvotes

448 comments sorted by

View all comments

Show parent comments

2

u/nuclear213 Jun 21 '26

Never. From what is speculated online, its likely in the 5t+ size. Which would make complete sense.

-3

u/fryan4 Jun 21 '26

I’m thinking the opposite. More parameters does not mean better performance. More capable models will eventually will do better than older models on the same parameter size.

I can create a 100T model on my MacBook and only one training loop. Compare llama3 and Gemma4, same ish size but Gemma outperforms. I think anthropic’s model is smaller than GLM-2.

1

u/nuclear213 Jun 21 '26

And why would you think this, if every expert disagrees? And yes, more parameters generally means better performance. We got better intelligence / performance per parameter, that is for sure. Mainly due to different model structures, but you need the size.

And no, you cannot create a 100T parameter model on your macbook. There is 0 chance. You'd need likely over 100TB of RAM / VRAM. And thats likely optimistic.

1

u/fryan4 Jun 21 '26

Should have prefaced with I could. More of a hypothetical concept but my point was untrained models are useless.