1 paper
Chushan Zhang, Ruihan Lu, Jinguang Tong +2
Leveraging 3D information within Multimodal Large Language Models (MLLMs) has recently shown significant advantages for indoor scene understanding. However, existing methods, inclu…