3 papers
cs.CL2026
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Meijia Chen, Hao Li, Zheng Lu +12
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces…
cs.SE2026
Zero2Repo: Can Coding Agents Build Repositories from Scratch?
Pei Yang, Tianyu Shi, Yuhang Yao +23
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and dep…
cs.LG2026
UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
Wenbo Zhang, Pengcheng Xu, Weizhi Du +2
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redi…