4 papers
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis
Long Zhang, Yuhan Chen, Chaoran Zhang +7
Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental is…
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Weikai Xu, Yunren Feng, Haoxiang Lei +6
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction polici…
CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning
Yuxuan Liu, Weikai Xu, Kun Huang +9
Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action funct…
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
Kun Huang, Weikai Xu, Yuxuan Liu +6
The Chain of Action-Planning Thoughts (CoaT) paradigm has been shown to improve the reasoning performance of VLM-based mobile agents in GUI tasks. However, the scarcity of diverse…