1 paper
Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10
Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…