benchmark 1computer-use agents 1cross-platform evaluation 1reward modeling 1vision-language models 1
From the 1 of 5 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex Team, B. An, B. Li +73
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, toge…
cs.AI2026
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
cs.AI2026
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
Zehao Li, Zhenyu Wu, Yibo Zhao +11
Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet training is highly sensitive to the quality of the reward func…