1 paper · 1 filter
Manuel Cherep, Pranav M R, Pattie Maes +1
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisio…