1 paper
Peng Ling, Yingda Yin, Lingting Zhu +5
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational…