1 paper
Quan Cui, Boyan Zhou, Yu Guo +4
Pioneering dual-encoder pre-training works (e.g., CLIP and ALIGN) have revealed the potential of aligning multi-modal representations with contrastive learning. However, these work…