1 paper · 1 filter
Rui Xie, Zhi Gao, Chenrui Shi +3
Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction. However, due to insufficient exposure to domain-s…