Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation
Anhao Zhao, Junlong Tong, Yingqi Fan +3
Standard on-policy distillation (OPD) for large language models estimates the reverse-KL objective using student-sampled tokens, yielding an unbiased single-sample Monte Carlo esti…
cs.LG2026
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
Anhao Zhao, Haoran Xin, Yingqi Fan +3
Knowledge distillation is central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (RL). We show that the prevailin…
cs.LG2025
The Few Govern the Many:Unveiling Few-Layer Dominance for Time Series Models
Xin Qiu, Junlong Tong, Yirong Sun +2
Large-scale models are at the forefront of time series (TS) forecasting, dominated by two paradigms: fine-tuning text-based Large Language Models (LLM4TS) and training Time Series…