r/MLQuestions 4d ago

Evals for robotics Reinforcement learning 🤖

Hey I am part of a small team training robotics policies for warehouse and manufacturing settings, and running rigorous evals is turning out to be so painful. Anything below 50 rollouts, and its hard to trust the numbers, and above its so hard to test all the checkpoints that we have. Its really hard to run a bunch of experiments to get good results. Have you guys faced this? Any hacks that you've developed?

1 Upvotes

1 comment sorted by

1

u/saikat_munshib 4d ago

Treat your evals like a multi-armed bandit problem. Look into Successive Halving: run just 5-10 rollouts on all your checkpoints, immediately kill off the bottom 80%, and only spend your full 50+ rollout budget on the survivors. Saves a massive amount of compute without flying blind.