4 papers
Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning
Tiehua Mei, Minxuan Lv, Leiyu Pan +5
Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally, potentially reinforcing flawed traces th…
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
Hongru Hou, Tiehua Mei, Denghui Geng +5
Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinforcement learning (RL) provide…
Heterogeneous Influence Maximization in User Recommendation
Hongru Hou, Jiachen Sun, Wenqing Lin +3
User recommendation systems enhance user engagement by encouraging users to act as inviters to interact with other users (invitees), potentially fostering information propagation.…
Predicting the critical behavior of complex dynamic systems via learning the governing mechanisms
Xiangrong Wang, Dan Lu, Zongze Wu +4
Critical points separate distinct dynamical regimes of complex systems, often delimiting functional or macroscopic phases in which the system operates. However, the long-term predi…