#computer-use agents
4 papers match
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
Woongkyu Lee, Jungwook Choi
The paper empirically studies how different inference-time scaling strategies affect the performance and failure modes of locally deployed autonomous computer-use agents under hard…
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Echoverse introduces a pipeline that compiles specifications into deep, stateful synthetic applications for training computer-use agents, using a co‑evolution loop that repairs env…
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.