1 paper
Shota Sato, Hajime Kiyama, Tosho Hirasawa +1
Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average…