1 paper
Alexandra Yakovleva, Henrik Pärssinen, Harri Valpola +2
Recent advances in vision-language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independent agents that reason and act…