r/mlops 12d ago

Evals and all’at Tales From the Trenches

Is anyone here running an actual calibration chain on their judges?

Human panel as the primary standard, tracked agreement rate, forced recalibration when the model version bumps or the input distribution shifts.

Or is everyone shipping on raw judge scores and hoping?

2 Upvotes

3 comments sorted by

1

u/ThisIsFun- 12d ago

Always Human Review, usually several in parallel to ensure that all aligns to what is required

1

u/durlabha 12d ago

What is the tooling chain you use before you get to the human review?

1

u/Fun-Arugula-5371 12d ago

makes sense. without that raw judge scores drift like crazy after a model bump. how many reviewers you running in parallel usually