Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs
Xixiang He, Xingming Li, Baiqi Wu +4
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable m…
cs.LG2026
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
Xixiang He, Qiyao Sun, Ao Cheng +5
Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improvin…