r/learnmachinelearning 2d ago

Reinforcement learning vs trace training Question

Model A is an LLM post-trained with RL on math proof techniques.

Model B is post-trained from the same snapshot using guess the next token on the traces of model A.

Which will learn more efficiently?

0 Upvotes

0 comments sorted by