1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Yuxi Chen, Haoyu Zhai, Chenkai Wang +4
GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots and directly interact with digital dev…