1 paper · 1 filter
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…