1 paper
Xianjin Wu, Dingkang Liang, Tianrui Feng +5
While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and…