Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Zhilin Huang, Hang Gao, Ziqiang Dong +6
Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes…
cs.LG2026
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
Feng Zhang, Xinhong Ma, Ziqiang Dong +5
Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…
cs.LG2025
Uncertainty-aware Reward Design Process
Yang Yang, Xiaolu Zhou, Bosong Ding +1
Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging process due to the inefficiencies and inconsistencies inherent in…