1 paper
Ziang Yan, Zhilin Li, Yinan He +9
Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals although they give comprehensive perception and reasoning in a spect…