1 paper
Jian Lan, Zhicheng Liu, Xinpeng Wang +5
The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related spatial intelligence tasks suc…