1 paper
Zuhui Wang, Yunting Yin, I. V. Ramakrishnan
Image-text matching aims to find matched cross-modal pairs accurately. While current methods often rely on projecting cross-modal features into a common embedding space, they frequ…