1 paper
Jin Xu, Xiaojian Huang, Zhuodong Luo +6
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the…