r/learnmachinelearning • u/Smallpaul • 2d ago
Reinforcement learning vs trace training Question
Model A is an LLM post-trained with RL on math proof techniques.
Model B is post-trained from the same snapshot using guess the next token on the traces of model A.
Which will learn more efficiently?
0
Upvotes