6 papers
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
Xinjie Chen, Biao Fu, Jing Wu +4
Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefull…
Entire Space Counterfactual Learning for Reliable Content Recommendations
Hao Wang, Zhichao Chen, Zhaoran Liu +4
Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias.…
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
Hao Wang, Zhiyu Wang, Yunlong Niu +5
Trustworthy process monitoring seeks to build an accurate and interpretable monitoring framework, which is critical for ensuring the safety of energy conversion plant (ECP) that op…
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
Xinjie Chen, Minpeng Liao, Guoxin Chen +4
Reinforcement learning with verifiable rewards (RLVR) has recently advanced the reasoning capabilities of large language models (LLMs). While prior work has emphasized algorithmic…
FreDF: Learning to Forecast in the Frequency Domain
Hao Wang, Licheng Pan, Zhichao Chen +6
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation…