1 paper · 1 filter
Hugues Thomas, Chen Chen, Jian Zhang
Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use vari…