4 papers
UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
Changxin Huang, Lv Tang, Zhaohuan Zhan +5
Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging.…
Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators
Qiwei Liang, Boyang Cai, Rongyi He +5
Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, exis…
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
Changxin Huang, Yanbin Chang, Junfan Lin +3
The ability to autonomously explore and resolve tasks with minimal human guidance is crucial for the self-development of embodied intelligence. Although reinforcement learning meth…
Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning
Runhao Zeng, Dingjie Zhou, Qiwei Liang +6
Learning behavior in legged robots presents a significant challenge due to its inherent instability and complex constraints. Recent research has proposed the use of a large languag…