1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Xinhan Zheng, Huyu Wu, Xueting Wang +2
Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from…