computer-use benchmarking 1interactive agents 1long-horizon tasks 1safety auditing 1tool-use evaluation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…
cs.LG2026
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
Sanjari Srivastava, Gang Li, Cheng Chang +8
Training web agents to navigate complex, real-world websites requires them to master - short-horizon interactions on multiple UI components (e.g., choosing the…