1 paper · 1 filter
Emmanuelle Bourigault
Vision-language models (VLMs) are increasingly used to answer questions about physical scenes, yet most evaluations reduce performance to a final answer. This hides whether the mod…