126 citations · 214 across the 16 of their papers we have counts for
1 paper · 1 filter
Yinpeng Dong, Huanran Chen, Jiawei Chen +6
Multimodal Large Language Models (MLLMs) that integrate text and other modalities (especially vision) have achieved unprecedented performance in various multimodal tasks. However,…