1 paper · 1 filter
Longxi Gao, Li Zhang, Pengzhi Gao +3
Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We in…