1 paper · 1 filter
Wenyao Zhang, Bozhou Zhang, Zekun Qi +3
Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image forecasting and 3D action prediction…