1 paper
Ruilin Yao, Shegnwu Xiong, Tianyu Zou +2
Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance deg…