3 papers
cs.AI2026
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Tianchen Guan, Xinlei Lin, Royce Cheng-Yue +2
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realisti…
cs.RO2026
PhysGraph: A Physics-aware 3D Scene Graph for Perception and Reasoning
Haoyu Li, Aaron Thomas, Shuyan Zhou +1
To perform a wide range of daily tasks, robots need to construct a 3D representation that is semantically rich, physically grounded, and structured enough to support task planning…
cs.CL2025
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
Yueqi Song, Ketan Ramaneti, Zaid Sheikh +18
Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this wo…