1 paper
Xilong Wang, John Bloch, Zedian Shao +3
Multi-modal large language model (MLLM)-based web agents interact with webpage environments by generating actions based on screenshots of the webpages. In this work, we propose Web…