collaborators

6 papers

cs.AI2026

JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR

Xinjie Chen, Biao Fu, Jing Wu +4

Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefull…

cs.LG2026

Entire Space Counterfactual Learning for Reliable Content Recommendations

Hao Wang, Zhichao Chen, Zhaoran Liu +4

Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias.…

cs.AI2026

An Accurate and Interpretable Framework for Trustworthy Process Monitoring

Hao Wang, Zhiyu Wang, Yunlong Niu +5

Trustworthy process monitoring seeks to build an accurate and interpretable monitoring framework, which is critical for ensuring the safety of energy conversion plant (ECP) that op…

cs.LG2026

CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks

Hao Wang, Licheng Pan, Zhichao Chen +7

Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…

cs.LG2025

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization

Xinjie Chen, Minpeng Liao, Guoxin Chen +4

Reinforcement learning with verifiable rewards (RLVR) has recently advanced the reasoning capabilities of large language models (LLMs). While prior work has emphasized algorithmic…

cs.LG2025

FreDF: Learning to Forecast in the Frequency Domain

Hao Wang, Licheng Pan, Zhichao Chen +6

Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation…