1 paper
Xingyu Ding, Yuzhong Zhao, Chunhai Zhao +3
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with lo…