activity
20242026
collaborators

20 papers

cs.AI2026

Task-Adaptive Rubrics for GUI Reward Modeling

Tao Xiong, Xavier Hu, Wenkai Wang +6

Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success cri…

cs.AI2026

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Long Zhang, Yuhan Chen, Chaoran Zhang +7

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental is…

cs.AI2026

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +29

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…

cs.AI2026

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

Guohong Liu, Jialei Ye, Pengzhi Gao +4

GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in…

cs.AI2026

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

Heng Qu, Yike Liu, Renren Jin +4

Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based…

cs.AI2026

How Mobile World Model Guides GUI Agents?

Weikai Xu, Kun Huang, Yunren Feng +10

Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences…