Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Tianchen Guan, Xinlei Lin, Royce Cheng-Yue +2
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realisti…
cs.AI2024★ 8 cited
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Tianbao Xie, Danyang Zhang, Jixuan Chen +14
Autonomous agents that accomplish complex computer tasks with minimal human interventions have the potential to transform human-computer interaction, significantly enhancing access…