2 papers
cs.LG2026
-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
Xianwei Chen, Shimin Zhang, Jibin Wu
Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency, but structurally d…
cs.LG2026
ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
Shimin Zhang, Xianwei Chen, Yufan Shen +2
Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated remarkable potential in enhancing the reasoning capability of Large Reasoning Models (LRMs). However…