1 paper
Federico Tavella, Amber Drinkwater, Angelo Cangelosi
Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment. In this work, we systematically evalu…