1 paper
Rodela Ghosh, Aviral Gupta, Guangjing Wang
Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We stud…