3 papers
cs.AI2026
Beyond Syntax: Action Semantics Learning for App Agents
Bohan Tang, Dezhao Luo, Jianheng Liu +5
The recent development of Large Language Models (LLMs) enables the rise of App agents that interpret user intent and operate smartphone Apps through actions such as clicking and sc…
cs.HC2025
ViMo: A Generative Visual GUI World Model for App Agents
Dezhao Luo, Bohan Tang, Kang Li +6
App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle w…
cs.CV2024
Motion Graph Unleashed: A Novel Approach to Video Prediction
Yiqi Zhong, Luming Liang, Bohan Tang +2
We introduce motion graph, a novel approach to the video prediction problem, which predicts future video frames from limited past data. The motion graph transforms patches of video…