2 papers
cs.LG2026
DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
Yuchen Xia, Qianguo Sun, Chao Song +3
While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead…
cs.LG2026
Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs
Chi Wang, Hanwen Wang, Yu Xia +2
We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision…