5 papers
Audio Spatially-Guided Fusion for Audio-Visual Navigation
Xinyu Zhou, Yinfeng Yu
Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achievi…
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
Angen Ye, Boyuan Wang, Chaojun Ni +21
World-Action Models (WAM) initialized from pre-trained video generation backbones have demonstrated remarkable potential for robot policy learning. However, existing approaches fac…
The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
Ziyu Wang, Chenyuan Liu, Yushun Xiang +16
Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack…
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
Xianchao Zeng, Xinyu Zhou, Youcheng Li +5
Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic manipulation, yet they remain limited in failure diagnosis and learning from failures. Add…
L1 Sample Flow for Efficient Visuomotor Learning
Weixi Song, Zhetao Chen, Tao Xu +6
Denoising-based models, such as diffusion and flow matching, have been a critical component of robotic manipulation for their strong distribution-fitting and scaling capacity. Conc…