3 papers
cs.AI2026
ExecCritic: Learn to Test, Test to Improve for Coding Agents
Leitian Tao, Baolin Peng, Haorui Wang +7
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode…
cs.SE2026
The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior
Xiangzhe Xu, Hamidreza Saghir, Qianhui Wu +5
As large language models continue to improve, agentic systems are becoming increasingly important, and tools are a key design dimension because they determine how agents access inf…
cs.CL2026
The Tool Illusion: Rethinking Tool Use in Web Agents
Renze Lou, Baolin Peng, Wenlin Yao +5
As web agents rapidly evolve, an increasing body of work has moved beyond conventional atomic browser interactions and explored tool use as a higher-level action paradigm. Although…