activity
20232026
most citedThe Oscars of AI Theater: A Survey on Role-Playing with Language Models

3 citations · 8 across the 31 of their papers we have counts for

collaborators
Showing 2026 · cs.AIShow all

6 papers · 2 filters

cs.AI2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the…

cs.AI2026

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

Jiawen Tao, Miao Peng, Yaoming Li +7

Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a…

cs.AI2026

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Qifan Zhang, Dongyang Ma, Tianqing Fang +5

Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human gui…

cs.AI2026

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Kanzhi Cheng, Zehao Li, Zheng Ma +11

Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap…

cs.AI2026

Exposing Weaknesses of Large Reasoning Models through Graph Algorithm Problems

Qifan Zhang, Jianhao Ruan, Aochuan Chen +4

Large Reasoning Models (LRMs) have advanced rapidly; however, existing benchmarks in mathematics, code, and common-sense reasoning remain limited. They lack long-context evaluation…

cs.AI2026

AlgBench: To What Extent Do Large Reasoning Models Understand Algorithms?

Henan Sun, Kaichi Yu, Yuyao Wang +5

Reasoning ability has become a central focus in the advancement of Large Reasoning Models (LRMs). Although notable progress has been achieved on several reasoning benchmarks such a…