Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Stopping and Routing LLM Judge Panels
Bin Zhu, Yi Xie, Yanghui Rao
LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The…
cs.CL2026
A Finite-Calibration Regime Map for LLM Judge Panels
Bin Zhu, Yi Xie, Yanghui Rao
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels…