From the 1 of 11 linked papers with an AI index.
4 papers · 1 filter
Simple-OPD: Demystifying Warm-up for On-policy Distillation
Tao Liu, Taiqiang Wu, Mao Zheng +5
On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage b…
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
Xuewei Yang, Jiachen Yu, Jie Wu +3
Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated…
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
Jiachen Yu, Zhihao Xu, Junjie Wang +1
Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. Howeve…
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
Tao Liu, Taiqiang Wu, Runming Yang +3
Supervised fine-tuning (SFT) is a fundamental post-training strategy to align Large Language Models (LLMs) with human intent. However, traditional SFT often ignores the one-to-many…