3 papers
cs.CL2026
Stopping and Routing LLM Judge Panels
Bin Zhu, Yi Xie, Yanghui Rao
LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The…
cs.AI2026
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning
Wenhao Zhang, Yibo Xie, Rui Wang +9
Autoregressive rollout generation is a major computational cost in reinforcement learning for large language models. Reusing each rollout batch for additional learner updates amort…
cs.CL2026
A Finite-Calibration Regime Map for LLM Judge Panels
Bin Zhu, Yi Xie, Yanghui Rao
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels…