1 paper
Yinfeng Wang, Zhiyuan Yao, Zheren Fu +2
Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, w…