2 papers
cs.CL2026
Self-Improving Large Language Models via Progressive Experience Evolution
Shijie Ren, Xiting Wang, Meng Li +8
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction expe…
cs.SE2026
Are Production Cloud Skills Adequately Tested? Measuring and Governing Skill Test Adequacy in Practice
Haotian Si, Junyi Chen, Shuyang Yu +5
Cloud platforms increasingly deliver reusable Cloud Skills that guide AI agents through multi-step resource operations, user choices, validation, and recovery. Existing Skill evalu…