agent evaluation 1agent harnesses 1behavior localization 1code analysis 1dense rewards 1llm-assisted tooling 1long-horizon planning 1multi-step tasks 1progressive disclosure 1software engineering 1terminal benchmarks 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Ruhan Wang, Yucheng Shi, Zongxia Li +7
The paper presents the Harness Handbook, a tool that automatically creates a behavior‑centric view of AI agent harness code using static analysis and LLM assistance, enabling devel…
cs.AI2026
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi +10
The paper presents Long-Horizon-Terminal-Bench, a benchmark of 46 extended tasks with fine-grained intermediate rewards to evaluate AI agents' long-horizon planning and debugging a…