15 papers
GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training
Zhihai Bi, Qiang Zhang, Guoyang Zhao +8
Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interacti…
GeoSem-WAM: Geometry- and Semantic-Aware World Action Models
Fulong Ma, Daojie Peng, Wenjun Yue +4
Recent World Action Models (WAMs) have demonstrated impressive capabilities in embodied decision-making. However, whether their effectiveness stems from explicit future imagination…
AttenA+: Rectifying Action Inequality in Robotic Foundation Models
Daojie Peng, Fulong Ma, Jiahang Cao +7
Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimizatio…
MoSA: Motion-constrained Stress Adaptation for Mitigating Real-to-Sim Gap in Continuum Dynamics via Learning Residual Anisotropy
Jiaxu Wang, Junhao He, Jingkai Sun +5
Learning real-world dynamics from visual observations is crucial for various domains. A common strategy is to calibrate simulators by estimating physical parameters, yet accuracy i…
Learning Structural Latent Points for Efficient Visual Representations in Robotic Manipulation
Yicheng Jiang, Jiaxu Wang, Junhao He +8
Current 3D-aware pretraining methods for embodied perception and manipulation are largely built on differentiable rendering frameworks, producing either fully implicit neural field…
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
Lingfeng Zhang, Xiaoshuai Hao, Qinwen Xu +7
Vision-and-language navigation (VLN) is a key task in Embodied AI, requiring agents to navigate diverse and unseen environments while following natural language instructions. Tradi…