8 papers
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Brian Wang, Bin Feng, Xiaoman Pan +26
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objec…
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
Weijia Zhang, Kunlun Zhu, Zeyi Liu +8
Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedback, yet their failures remain difficult to…
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
Kunlun Zhu, Xuyan Ye, Zhiguang Han +9
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but pro…
VISTA: View-Consistent Self-Verified Training for GUI Grounding
Xinyu Qiu, Yunzhu Zhang, Heng Jia +3
When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult…
Foundation Protocol: A Coordination Layer for Agentic Society
Bang Liu, Yongfeng Gu, Jiayi Zhang +26
Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly interact with one another. A…
A Rubric-Supervised Critic from Sparse Real-World Outcomes
Xingyao Wang, Valerie Chen, Heng Ji +1
Academic benchmarks for coding agents tend to reward autonomous task completion, measured by verifiable rewards such as unit-test success. In contrast, real-world coding agents ope…