1 paper · 1 filter
Yixuan Li, Yuhui Chen, Mingcai Zhou +3
Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existing approaches often lack the abil…