1 paper
Liuyang Song, Yi Zhang, Zhongyi Deng +2
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are a…