2 papers
cs.AI2026
Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
Yongkang Yang, Zhezheng Hao, Hong Zhang +8
On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research…
cs.AI2026
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
Ziyan Liu, Zhezheng Hao, Yeqiu Chen +7
Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train…