activity
20242026
most citedOS-Copilot: Towards Generalist Computer Agents with Self-Improvement

5 citations · 10 across the 22 of their papers we have counts for

collaborators
Showing 2026Show all

12 papers · 1 filter

cs.AI2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the…

cs.AI2026

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Kai Chen, Zichen Ding, Jiaye Ge +20

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly…

cs.AI2026

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

Yikun Fu, Bowen Fu, Zhenyu Wu +10

Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, ex…

cs.CL2026

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

Jianxiang Yu, Jiapeng Zhu, Bochen Lin +3

LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon interactive tasks. Existing ski…

cs.CL2026

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

Yibo Zhao, Zichen Ding, Jiayi Wu +2

Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reasoning. However, the rapid gro…

cs.AI2026

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Kanzhi Cheng, Zehao Li, Zheng Ma +11

Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap…