1 paper
Davide Berasi, Matteo Farina, Massimiliano Mancini +2
Vision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While prior works demonstrated that VLMs…