1 paper
Kang He, Yuzhe Ding, Haining Wang +3
Previous multimodal sentence representation learning methods have achieved impressive performance. However, most approaches focus on aligning images and text at a coarse level, fac…