2 papers
cs.LG2026
Calibration-Aware Policy Optimization for Reasoning LLMs
Ziqi Wang, Xingzhou Lou, Meiqi Wu +2
Group Relative Policy Optimization (GRPO) enhances LLM reasoning but often induces overconfidence, where incorrect responses yield lower perplexity than correct ones, degrading rel…
cs.LG2025
Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Xingzhou Lou, Dong Yan, Wei Shen +3
Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of…