3 papers
cs.LG2026
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
Zhiwei Zhang, Zechen Sun, Fei Zhao +6
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supe…
cs.AI2026
Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning
Zilin Zhao, Han Yang, Tianpei Yang +7
Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and…
cs.CL2026
Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback
Kang Peng, Zhiwei Zhang, Yichen Zhang +7
Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following proc…