1 paper
Yilian Liu, Sicong Leng, Guoshun Nan +7
Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffect…