From the 2 of 23 linked papers with an AI index.
23 papers
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Yikun Fu, Bowen Fu, Zhenyu Wu +10
Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, ex…
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
Hang Yan, Fangzhi Xu, Qiushi Sun +14
The rapid advancement of Large Language Models (LLMs) has catalyzed the development of autonomous agents capable of navigating complex environments. However, existing evaluations p…
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
Jianxiang Yu, Jiapeng Zhu, Bochen Lin +3
LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon interactive tasks. Existing ski…
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
Yibo Zhao, Zichen Ding, Jiayi Wu +2
Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reasoning. However, the rapid gro…