1 paper
Jingxuan Wei, Xi Bai, Shan Liu +8
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a f…