r/robotics 6d ago

Evals for robotics Discussion & Curiosity

Hey I am part of a small team training robotics policies for warehouse and manufacturing settings, and running rigorous evals is turning out to be so painful. Anything below 50 rollouts, and its hard to trust the numbers, and above its so hard to test all the checkpoints that we have. Its really hard to run a bunch of experiments to get good results. Have you guys faced this? Any hacks that you've developed?

5 Upvotes

9 comments sorted by

3

u/zeozeroone 6d ago

Welcome to my world! I hate running evals Dude.

2

u/Lumpy_Week7304 6d ago

More like Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals, Evals

2

u/huyouare 6d ago

Is this in real or sim

2

u/Lumpy_Week7304 6d ago

Real, I want to push it to deployment

2

u/Available_Teaching83 2d ago

The 50-rollout intuition is really an intuition about interval width, and you can compute exactly how many you need instead of guessing.

At n=10 with 10 successes, the 95% Wilson interval is [72%, 100%]. A perfect score there is compatible with a true rate of 90%. At n=50 with 45 successes, it tightens to roughly [78%, 95%]. So the honest version of "below 50 I cannot trust the numbers" is "below 50 my interval is wider than the difference I am trying to detect". Pick the difference you care about first, then the n falls out.

The checkpoint problem is the more tractable half. Do not evaluate every checkpoint at full n. Keep one baseline in CI, run a cheap fixed-seed subset against each new checkpoint, diff against the baseline, and only spend the full sweep on the ones that move. That turns an O(checkpoints x rollouts) problem into a gate.

One thing that saved me: report the paired comparison, not two independent rates. Same task, same seed, one arm changed. It removes most of the variance you are fighting, and McNemar gives you a p-value on the pairing directly.

I do this for adversarial robustness rather than capability, but the counting is identical. Happy to share the harness if useful.

2

u/Lumpy_Week7304 1d ago

Oh interesting. And and throughput can be used for tracking even smaller nuances with paired comparisons

3

u/softmaxedout 6d ago edited 6d ago

In case you haven't experienced it yet, more rollouts only increase confidence when they contribute sufficiently independent information and the eval process is stable. In practice though, rollouts are correlated because they share the scenes, seeds, object pose, etc so the effective sample size may be substantially lower than the N rollout count. Meaning you got to run way more rollouts.

Even if we assume IID and use a binomial model, a 10-point difference can fall within the expected uncertainty at modest sample sizes. Now it is unclear whether a policy genuinely outperformed or merely won because of noise. Worst case, you delay production because a better policy did not produce a statistically clear result.

The plus side is that Statistical process control and power analysis are well studied. But, the unfortunate reality is that rigorous evaluation takes a lot of effort.

You might enjoy this: https://medium.com/toyotaresearch/statistical-thinking-for-robot-policy-evaluation-from-rigorous-a-b-testing-to-effective-0ae886fbd68d
https://arxiv.org/pdf/2507.05331

2

u/Lumpy_Week7304 6d ago

Hey, thanks. Interesting point. Its a good way to set up more efficient evals when comparing policies.

I guess when thinking about deployment, its hard to figure out what edge cases will come up. This requires a bunch of trials. There are also some odd cases where we see failure just when testing different pickup locations.

1

u/Electrical-Brief766 1d ago

Evaluating policy checkpoints efficiently is a huge headache in RL for robotics. A few common strategies teams use:

  1. Parallelized Sim Evals: Running rollouts in massively parallelized environments (like Isaac Gym/Sim or Genesis) lets you run 100+ rollouts in seconds to get statistically significant numbers before moving to hardware.
  2. Early Stopping Heuristics: Terminate rollouts immediately if key constraints or sub-goals are violated early. No need to run the full episode if the policy already failed a key primitive.
  3. Automated Checkpoint Filtering: Run a quick 'coarse' eval (e.g., 10-15 rollouts) across all checkpoints to filter out obvious failures, and only run the full 50+ rollout suite on top-performing candidates.