1 paper
Qianqi Yan, Yue Fan, Hongquan Li +5
Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inc…