1 paper · 1 filter
Revanth Gangi Reddy, Sagnik Mukherjee, Jeonghwan Kim +3
Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of…