1 paper · 1 filter
WenZhang Wei, Zhipeng Gui, Dehua Peng +2
The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classifica…