1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2026
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
Li Wang, Xiaodong Lu, Xiaohan Wang +4
Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already pres…
cs.LG2026★ 1 cited
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…