3 papers
cs.LG2026
Distillation as Probability Transport: Routed On-Policy Distillation
Tianle Xia, Lingxiang Hu, Yiding Sun +6
On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on ind…
cs.AI2026
When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
Linfang Shang, Ming Xu, Yiding Sun +4
Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, ne…
cs.CL2026
GraSP: Graph-Structured Skill Compositions for LLM Agents
Tianle Xia, Lingxiang Hu, Yiding Sun +5
Skill ecosystems for LLM agents have matured rapidly, yet recent benchmarks show that providing agents with more skills does not monotonically improve performance -- focused sets o…