5 citations · 5 across the 1 of their papers we have counts for
1 paper
Jiali Duan, Liqun Chen, Son Tran +4
Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusi…