1 paper
Rhea Chowers, Oshri Naparstek, Udi Barzelay +1
Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existing models show a strong modality…