Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
Siqi Zhu, Xuyan Ye, Hongyu Lu +2
On-policy distillation (OPD) and on-policy self-distillation (OPSD) have emerged as promising post-training methods for large language models, offering dense token-level supervisio…
cs.AI2025
Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
Yihong Liu, Junyi Li, Wayne Xin Zhao +2
Combinatorial optimization problems are traditionally tackled with handcrafted heuristic algorithms, which demand extensive domain expertise and significant implementation effort.…