2 papers
cs.RO2026
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
Ziyao Wang, Bingying Wang, Hanrong Zhang +7
Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this…
cs.CV2026
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
Guoheng Sun, Tingting Du, Kaixi Feng +6
Vision-Language-Action (VLA) models enable instruction-following robotic manipulation, but they are typically pretrained on 2D data and lack 3D spatial understanding. An effective…