8 papers
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
Huosen Ou, Dongni Song, Yuncong Wang +2
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabul…
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
Yihao Qin, Yuanfei Wang, Hang Zhou +3
Decision transformer based sequential policies have emerged as a powerful paradigm in offline reinforcement learning (RL), yet their efficacy remains constrained by the quality of…
MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction
Qiang Zhang, Jiahao Ma, Peiran Liu +20
Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-l…
Physics-informed Diffusion Mamba Transformer for Real-world Driving
Hang Zhou, Qiang Zhang, Peiran Liu +3
Autonomous driving systems demand trajectory planners that not only model the inherent uncertainty of future motions but also respect complex temporal dependencies and underlying p…
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
Yunheng Wang, Yuetong Fang, Taowen Wang +6
Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied…
TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation
Peiran Liu, Qiang Zhang, Daojie Peng +6
Object Navigation (ObjectNav) has made great progress with large language models (LLMs), but still faces challenges in memory management, especially in long-horizon tasks and dynam…