3 papers
cs.LG2026
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
Yubo Zhang, Xinhong Ma, Zezhong Tan +1
Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Pr…
cs.LG2026
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Zhilin Huang, Hang Gao, Ziqiang Dong +6
Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes…
cs.AI2025
Towards Flash Thinking via Decoupled Advantage Policy Optimization
Zezhong Tan, Hang Gao, Xinhong Ma +2
Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although exi…