1 paper
Jiangye Yuan, Gowri Kumar, Baoyuan Wang
While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this…