1 paper · 1 filter
Lan Wei, Kangyi Lu, Yongchen Wang +5
Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliable execution still hinges on identifying which physical object an inst…