8 papers
On Safety Risks in Experience-Driven Self-Evolving Agents
Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8
Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…
Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code
Yang Liu, Da Song, Armstrong Foundjem +2
Chain-of-Thought (CoT) prompting is widely used to elicit explicit reasoning from large language models for code (LLM4Code). However, its impact on robustness and the stability of…
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun +7
Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequence…
Intent-aligned Formal Specification Synthesis via Traceable Refinement
Zhe Ye, Aidan Z. H. Yang, Huangyuan Su +6
Large language models are increasingly used to generate code from natural language, but ensuring correctness remains challenging. Formal verification offers a principled way to obt…
CUBE: A Standard for Unifying Agent Benchmarks
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…
Self-Sovereign Agent
Wenjie Qu, Xuandong Zhao, Jiaheng Zhang +1
We investigate the emerging prospect of self-sovereign agents -- AI systems that can economically sustain and extend their own operation without human involvement. Recent advances…