entropy preservation 1large language models 1reinforcement learning 1self-distillation 1supervised fine-tuning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
Hao Wang, Hao Gu, Hongming Piao +6
The paper introduces CurioSFT, an entropy-preserving supervised fine-tuning approach that uses adaptive self-distillation to keep exploration abilities in large reasoning models, l…
cs.RO2025
Real-world Reinforcement Learning from Suboptimal Interventions
Yinuo Zhao, Huiqian Jin, Lechun Jiang +9
Real-world reinforcement learning (RL) offers a promising approach to training precise and dexterous robotic manipulation policies in an online manner, enabling robots to learn fro…