11 papers
ADAPT: Analytical Disturbance-Aware Policy Training for Humanoid Locomotion
Bofan Lyu, Jindou Jia, Kuangji Zuo +7
Humanoids deployed in human-centered environments must handle force-interactive tasks, where external contacts introduce unexpected disturbances that disrupt locomotion accuracy an…
APEX: Adaptive Policy Execution for Precise Manipulation
Mengfei Zhao, Chenxi Jiang, Tuo An +2
Modern imitation learning methods, including visuomotor and Vision-Language-Action (VLA) policies, typically output high-level action references that are executed by low-level cont…
GIVE: Grounding Human Gestures in Vision-Language-Action Models
Pengfei Liu, Gen Li, Junqiao Fan +4
Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action…
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Yikai Tang, Haoran Geng, Jindou Jia +5
Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy gen…
MARS Policy: Multimodality Only When It Matters
Jindou Jia, Tuo An, Yuxuan Hu +7
Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavior…
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
Jingliang Li, Jindou Jia, Tuo An +7
When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects ma…