1 paper · 1 filter
Yuhang Zang, Wei Li, Jun Han +2
Recent Multimodal Large Language Models (MLLMs) are remarkable in vision-language tasks, such as image captioning and question answering, but lack the essential perception ability,…