1 paper · 1 filter
Xiaowen Sun, Matthias Kerzel, Mengdi Li +3
Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle with precise spatial reasoning,…