1 paper
Bin Zhu, Yi Xie, Yanghui Rao
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels…