20 papers
Beyond Syntax: Action Semantics Learning for App Agents
Bohan Tang, Dezhao Luo, Jianheng Liu +5
The recent development of Large Language Models (LLMs) enables the rise of App agents that interpret user intent and operate smartphone Apps through actions such as clicking and sc…
Generative Models in Decision Making: A Survey
Xinyu Shao, Jianping Zhang, Haozhi Wang +9
Generative models have fundamentally reshaped the landscape of decision-making, reframing the problem from pure scalar reward maximization to high-fidelity trajectory generation an…
K^2-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control
Zhe Wu, Donglin Mo, Hongjin Lu +7
Existing mobile device control agents often perform poorly when solving complex tasks requiring long-horizon planning and precise operations, typically due to a lack of relevant ta…
ActionCodec: What Makes for Good Action Tokenizers
Zibin Dong, Yicheng Liu, Shiduo Zhang +8
Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training eff…
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
Zhe Wu, Hongjin Lu, Junliang Xing +10
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on dire…
Collab-Solver: Collaborative Solving Policy Learning for Mixed-Integer Linear Programming
Siyuan Li, Yifan Yu, Zhihao Zhang +5
Mixed-integer linear programming (MILP) has been a fundamental problem in combinatorial optimization. Conventional MILP solving mainly relies on carefully designed heuristics embed…