1 paper
Mingyeong Kim, Jungwon Choi, Chaeyun Jang +1
Vision-language models (VLMs) are often deployed on text-only inputs, although they are trained with images. We find that removing the vision modality causes large drops in accurac…