8 citations · 15 across the 3 of their papers we have counts for
1 paper · 1 filter
Zhihuan Jiang, Zhen Yang, Jinhao Chen +4
Multi-modal large language models (MLLMs) have demonstrated promising capabilities across various tasks by integrating textual and visual information to achieve visual understandin…