4 papers
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Wenpu Liu, Yuqi Xu, Weichu Xie +8
Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective…
Step-wise Rubric Rewards for LLM Reasoning
Weichu Xie, Haozhe Zhao, Wenpu Liu +15
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision ov…
User-Creator Feature Polarization in Recommender Systems with Dual Influence
Tao Lin, Kun Jin, Andrew Estornell +3
Recommender systems serve the dual purpose of presenting relevant content to users and helping content creators reach their target audience. The dual nature of these systems natura…
Toward Optimal LLM Alignments Using Two-Player Games
Rui Zheng, Hongyi Guo, Zhihan Liu +10
The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. Howev…