1 paper · 1 filter
Yueen Ma, Dafeng Chi, Shiguang Wu +3
Vision-language-action models have gained significant attention for their ability to model multimodal sequences in embodied instruction following tasks. However, most existing mode…