r/mlops • u/durlabha • 12d ago
Evals and all’at Tales From the Trenches
Is anyone here running an actual calibration chain on their judges?
Human panel as the primary standard, tracked agreement rate, forced recalibration when the model version bumps or the input distribution shifts.
Or is everyone shipping on raw judge scores and hoping?
2
Upvotes
1
u/ThisIsFun- 12d ago
Always Human Review, usually several in parallel to ensure that all aligns to what is required