2 papers
cs.AI2026
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of f…
cs.LG2025
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
Sanjari Srivastava, Gang Li, Cheng Chang +8
Training web agents to navigate complex, real-world websites requires them to master - short-horizon interactions on multiple UI components (e.g., choosing the…