14 papers
Darwin Mobile Agent: A Roadmap for Self-Evolution
Daniel Beechey, Derek Yuen, Jianheng Liu +5
The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments. Guided by the "Bitter Lesson", we argue that the most eff…
Beyond Syntax: Action Semantics Learning for App Agents
Bohan Tang, Dezhao Luo, Jianheng Liu +5
The recent development of Large Language Models (LLMs) enables the rise of App agents that interpret user intent and operate smartphone Apps through actions such as clicking and sc…
Generative Models in Decision Making: A Survey
Xinyu Shao, Jianping Zhang, Haozhi Wang +9
Generative models have fundamentally reshaped the landscape of decision-making, reframing the problem from pure scalar reward maximization to high-fidelity trajectory generation an…
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
Zhe Wu, Hongjin Lu, Junliang Xing +10
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on dire…
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
Haoyu Zhao, Weizhong Ding, Yuhao Yang +4
Recent advances in Multimodal Large Language Models (MLLMs) have enabled their use as intelligent agents for smartphone operation. However, existing methods depend on the Android D…
Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
Georgios Papoudakis, Thomas Coste, Jianye Hao +2
Reinforcement learning (RL) using foundation models for policy approximations in multi-turn tasks remains challenging. We identify two main limitations related to sparse reward set…