1 paper
Ruoqi Guo, Yi Liu, Gelei Deng +7
Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot re…