1 paper · 1 filter
Xin Wang, Zhiyao Cui, Hao Li +10
Vision language model (VLM)-based mobile agents show great potential for assisting users in performing instruction-driven tasks. However, these agents typically struggle with perso…