1 paper · 1 filter
Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto +2
Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image capt…