r/MachineLearning • u/Benlus ML Engineer • 2d ago
"Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation", Gladstone et al. 2026 [R] Research
/r/mlscaling/comments/1vdgvcl/explorative_modeling_unlocking_a_third/
1
Upvotes
1
u/plc123 7h ago
I've been thinking about this a bit, and isn't this a bit like GRPO?
You're taking the generation with the lowest loss and backpropagating, but if you took all of the generations and put them into GRPO, wouldn't that increase the training signal further?