1 paper
Zhongbin Guo, Jiahe Liu, Yushan Li +5
Existing Vision Language Models (VLMs) architecturally rooted in "flatland" perception, fundamentally struggle to comprehend real-world 3D spatial intelligence. This failure stems…