4 papers · 1 filter
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Wenpu Liu, Yuqi Xu, Weichu Xie +8
Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective…
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning
Ziyue Wang, Aomufei Yuan, Yongfu Zhu +10
Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce eac…
Step-wise Rubric Rewards for LLM Reasoning
Weichu Xie, Haozhe Zhao, Wenpu Liu +15
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision ov…
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Yuqi Xu, Rizhen Hu, Zihan Liu +2
Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searchi…