From the 1 of 8 linked papers with an AI index.
8 papers
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
Hang Yan, Fangzhi Xu, Qiushi Sun +14
The rapid advancement of Large Language Models (LLMs) has catalyzed the development of autonomous agents capable of navigating complex environments. However, existing evaluations p…
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
Kanzhi Cheng, Zehao Li, Zheng Ma +11
Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap…
TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents
Hang Yan, Xinyu Che, Fangzhi Xu +7
Recent advances in autonomous LLM agents demonstrate their ability to improve performance through iterative interaction with the environment. We define this paradigm as Test-Time I…
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent
Bowen Yang, Kaiming Jin, Zhenyu Wu +12
While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalizati…
OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
Zhenyu Wu, Jingjing Xie, Zehao Li +8
With VLM-powered computer-using agents (CUAs) becoming increasingly capable at graphical user interface (GUI) navigation and manipulation, reliable step-level decision-making has e…