1 paper · 1 filter
Qixiang Chen, Cheng Zhang, Chi-Wing Fu +2
Recent multimodal large language models (MLLMs) show great potential in natural image understanding. Yet, they perform well, mainly on reasoning in-view contents within the image f…