3 papers
cs.SE2026
From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws
Mengzhuo Chen, Junjie Wang, Zhe Liu +3
LLM agents increasingly rely on agent harness: the runtime infrastructure around the base model that defines execution environments, tool interfaces, context, lifecycle orchestrati…
cs.AI2026
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
Hongming Piao, Chi Liu, Mengzhuo Chen +5
Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence. The former enables autonomous retrieval a…
cs.MA2026
Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
Mengzhuo Chen, Junjie Wang, Fangwen Mu +4
Failure attribution, i.e., identifying the responsible agent and decisive step of a failure, is particularly challenging in LLM-based multi-agent systems (MAS) due to their natural…