1 paper
Yulin Chen, Tri Cao, Haoran Li +7
Web agents powered by vision-language models (VLMs) enable autonomous interaction with web environments by perceiving and acting on both visual and textual webpage content to accom…