1 paper · 1 filter
Shanglin Yuan, Weiheng Zhao, Xianda Guo +4
Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…