8 papers · 1 filter
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
Huosen Ou, Dongni Song, Yuncong Wang +2
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabul…
MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction
Qiang Zhang, Jiahao Ma, Peiran Liu +20
Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-l…
Physics-informed Diffusion Mamba Transformer for Real-world Driving
Hang Zhou, Qiang Zhang, Peiran Liu +3
Autonomous driving systems demand trajectory planners that not only model the inherent uncertainty of future motions but also respect complex temporal dependencies and underlying p…
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
Yunheng Wang, Yuetong Fang, Taowen Wang +6
Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied…
TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation
Peiran Liu, Qiang Zhang, Daojie Peng +6
Object Navigation (ObjectNav) has made great progress with large language models (LLMs), but still faces challenges in memory management, especially in long-horizon tasks and dynam…
A Value Function Space Approach for Hierarchical Planning with Signal Temporal Logic Tasks
Peiran Liu, Yiting He, Yihao Qin +2
Signal Temporal Logic (STL) has emerged as an expressive language for reasoning intricate planning objectives. However, existing STL-based methods often assume full observation and…