1 paper
Xiyin Zeng, Yuyu Sun, Haoyang Li +2
Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or…