2 papers
cs.CL2026
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Ziqi Zhao, Xinyu Ma, Liu Yang +6
On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, ex…
cs.CL2025
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
Jitai Hao, Qiang Huang, Hao Liu +3
Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key c…