1 paper
Yunqing Hu, Zheming Yang, Chang Zhao +4
Multimodal large language models (MLLMs) demonstrate exceptional capabilities in semantic understanding and visual reasoning, yet they still face challenges in precise object local…