1 paper
Zhe Wu, Hongjin Lu, Junliang Xing +10
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on dire…