3 papers
cs.SE2026
EnvPilot: Systematic Design and Evaluation of an Experience-Augmented Agent for Software Environment Setup
Hanwu Chen, Hanyu Lin, Zhanjiang Yang +6
Environment Setup is a critical yet complex task in software engineering that relies heavily on expert knowledge. Existing automated environment setup methods lack the ability to a…
cs.SE2026
Are Production Cloud Skills Adequately Tested? Measuring and Governing Skill Test Adequacy in Practice
Haotian Si, Junyi Chen, Shuyang Yu +5
Cloud platforms increasingly deliver reusable Cloud Skills that guide AI agents through multi-step resource operations, user choices, validation, and recovery. Existing Skill evalu…
cs.CL2026
Self-Improving Large Language Models via Progressive Experience Evolution
Shijie Ren, Xiting Wang, Meng Li +8
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction expe…