1 paper
Junwon You, Dasol Kang, Jae-Hun Jung
Contrastive Vision-Language Models (VLMs) have demonstrated strong zero-shot capabilities. However, their cross-modal alignment remains biased toward English due to limited multili…