99 citations · 137 across the 12 of their papers we have counts for
1 paper · 1 filter
Jinyu Yang, Jiali Duan, Son Tran +6
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…