6 papers
SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering
Jiujiu Chen, Yazheng Liu, Sihong Xie +1
Large language models excel at complex reasoning, yet evaluating their intermediate steps remains challenging. Although process reward models provide step-wise supervision, they of…
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
Yao Shu, Chenxing Wei, Hongbin Lin +2
Online reinforcement learning with verifiable rewards (RLVR) turns checkable outcomes into a scalable training signal, but it keeps rollout generation, verifier scoring, and refere…
Robust Conditional Conformal Prediction via Branched Normalizing Flow
Rui Xu, Xingyuan Chen, Wenxing Huang +4
Conformal prediction (CP) constructs prediction sets with marginal coverage guarantees under the assumption that the calibration and test distributions are identical. However, unde…
Geometry-Calibrated Conformal Abstention for Language Models
Rui Xu, Yi Chen, Sihong Xie +1
When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about…
GFM4GA: Graph Foundation Model for Group Anomaly Detection
Jiujiu Chen, Weijun Zeng, Shaofeng Hu +2
Group anomaly detection is crucial in many network applications, but faces challenges due to diverse anomaly patterns. Motivated by the success of large language models (LLMs) in n…
Perturbation-mitigated USV Navigation with Distributionally Robust Reinforcement Learning
Zhaofan Zhang, Minghao Yang, Sihong Xie +1
The robustness of Unmanned Surface Vehicles (USV) is crucial when facing unknown and complex marine environments, especially when heteroscedastic observational noise poses signific…