1 paper
Simone Alghisi, Gabriel Roccabruna, Massimo Rizzoli +2
Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. Howev…