3 papers
cs.AI2026
SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering
Jiujiu Chen, Yazheng Liu, Sihong Xie +1
Large language models excel at complex reasoning, yet evaluating their intermediate steps remains challenging. Although process reward models provide step-wise supervision, they of…
cs.LG2026
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
Yao Shu, Chenxing Wei, Hongbin Lin +2
Online reinforcement learning with verifiable rewards (RLVR) turns checkable outcomes into a scalable training signal, but it keeps rollout generation, verifier scoring, and refere…
cs.LG2026
Robust Conditional Conformal Prediction via Branched Normalizing Flow
Rui Xu, Xingyuan Chen, Wenxing Huang +4
Conformal prediction (CP) constructs prediction sets with marginal coverage guarantees under the assumption that the calibration and test distributions are identical. However, unde…