Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Variational Adapter for Cross-modal Similarity Representation
WenZhang Wei, Zhipeng Gui, Dehua Peng +2
The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classifica…
cs.CV2023
Dynamic Visual Semantic Sub-Embeddings and Fast Re-Ranking
Wenzhang Wei, Zhipeng Gui, Changguang Wu +3
The core of cross-modal matching is to accurately measure the similarity between different modalities in a unified representation space. However, compared to textual descriptions o…