Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
Yutong Chen, Jiandong Gao, Ji Wu
Training Large Reasoning Model (LRM) is usually unstable and unpredictable, especially on hard problems or weak foundation models. We found that the current post-training scaling s…
cs.LG2026
Collaborative Parameter Learning: Mitigating Forgetting via Parameter-Level Gradient Analysis
Mutian Yang, Zisen Zhan, Yutong Chen +7
Catastrophic forgetting during knowledge injection impairs the ability of large language models to acquire new knowledge without overwriting previously mastered knowledge. Recent s…
cs.LG2025
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
Yutong Chen, Jiandong Gao, Ji Wu
R1-style Reinforcement Learning (RL) significantly enhances Large Language Models' reasoning capabilities, yet the mechanism behind rule-based RL remains unclear. We found that sma…