1 paper
Yuheng Li, Haotian Liu, Mu Cai +5
In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language mo…