1 paper
Yuzhe Zhang, Xianwei Xue, Xingyong Wu +8
Autonomous GUI agents based on vision-language models (VLMs) often assume deterministic environment responses, generating actions without verifying whether previous operations succ…