4 papers
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
Kuangji Zuo, Gen Li, Bofan Lyu +9
Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insuff…
MARS Policy: Multimodality Only When It Matters
Jindou Jia, Tuo An, Yuxuan Hu +7
Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavior…
Feedback World Model Enables Precise Guidance of Diffusion Policy
Tuo An, Jindou Jia, Gen Li +8
World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encoun…
FLASH: Efficient Visuomotor Policy via Sparse Sampling
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference…