1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Xiang Ma, Lexin Fang, Litian Xu +1
Cross-modal alignment is a crucial task in multimodal learning aimed at achieving semantic consistency between vision and language. This requires that image-text pairs exhibit simi…