1 paper · 1 filter
Marco Mistretta, Alberto Baldrati, Lorenzo Agnolucci +2
Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individuall…