1 paper
Qixiang Chen, Cheng Zhang, Chi-Wing Fu +2
Recent multimodal large language models (MLLMs) show great potential in natural image understanding. Yet, they perform well, mainly on reasoning in-view contents within the image f…