3 papers
cs.AI2026
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
Siyuan Gan, Yuhan Li, Xiran Wang +6
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO.…
cs.AI2026
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
Siyuan Gan, Yuhan Li, Xiran Wang +5
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and cap…
cs.AI2026
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
Siyuan Gan, Jiaheng Liu, Boyan Wang +8
Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from thinking, a long Chain of Thought (Co…