1 paper · 1 filter
Guoheng Sun, Tingting Du, Kaixi Feng +6
Vision-Language-Action (VLA) models enable instruction-following robotic manipulation, but they are typically pretrained on 2D data and lack 3D spatial understanding. An effective…