1 paper
Yue Cao, Yangzhou Liu, Zhe Chen +4
Despite significant advancements in Multimodal Large Language Models (MLLMs) for understanding complex human intentions through cross-modal interactions, capturing intricate image…