2 papers
cs.AI2026
RobustSGPO: Search-Space Control for Agent Harness Evolution
Zibo Zhao, Jijun Shi, Mo Zhou +9
Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unreso…
cs.AI2026
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Yunfei Zhang, Boyu Feng, Changhua Pei +14
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…