5 citations · 5 across the 4 of their papers we have counts for
1 paper · 1 filter
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu +8
Existing efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are of…