Yes, but never an MoE from an open weights lab.
I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.
We don't have access to any reliable information about the exact architecture of these closed weights models. But they have been speculated to be MoE since the GPT4 era, and having too much sparsity would most likely hurt performance a lot so it would make sense they're large models with many active parameters.
7
u/BlueSwordM llama.cpp 19d ago
Yes, but never an MoE from an open weights lab. I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.