Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR
Chuanyu Qin, Chenxu Yang, Qingyi Si +3
Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…
cs.LG2026
Co-Evolving Policy Distillation
Naibin Gu, Chenxu Yang, Qingyi Si +7
RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…
cs.LG2026
Self-Distilled RLVR
Chenxu Yang, Chuanyu Qin, Qingyi Si +7
On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals…