1 paper · 1 filter
Yuhang Su, Mei Wang, Yaoyao Zhong +4
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature…