1 paper · 1 filter
Siyu Ma, Yuqi Liang, Chang Yu +5
Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state requ…