From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…
cs.AI2025
Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
Zihan Dong, Xinyu Fan, Zixiang Tang +1
Controlling desktop applications via software remains a fundamental yet under-served problem. Existing multi-modal large language models (MLLMs) ingest screenshots and task instruc…