1 paper · 1 filter
Xianzhe Fan, Shengliang Deng, Xiaoyang Wu +7
Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D informa…