2 papers
cs.LG2026
Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs
Chi Wang, Hanwen Wang, Yu Xia +2
We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision…
cs.LG2026
DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
Yuchen Xia, Qianguo Sun, Chao Song +3
On-policy distillation (OPD) trains student models on their own rollouts to reduce exposure bias. However, in multi-turn agent scenarios, early student errors can lead a trajectory…