1 paper · 1 filter
Chengwei Ma, Zhen Tian, Zhou Zhou +5
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the…