11 papers
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
Jingzhi Huang, Junkai Huang, Wenxuan Song +4
Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories t…
Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using a Feed-Forward 3D Model
Yuantai Zhang, Jiaqi Yang, Huajian Zeng +5
Fast and reliable initialization is critical for monocular visual-inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Des…
AERR-Nav: Adaptive Exploration-Recovery-Reminiscing Strategy for Zero-Shot Object Navigation
Jingzhi Huang, Junkai Huang, Haoyang Yang +2
Zero-Shot Object Navigation (ZSON) in unknown multi-floor environments presents a significant challenge. Recent methods, mostly based on semantic value greedy waypoint selection, s…
PNav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
Tianfu Li, Wenbo Chen, Haoxuan Xu +2
In Vision-and-Language Navigation (VLN), an agent is required to plan a path to the target specified by the language instruction, using its visual observations. Consequently, preva…
OmniDP: Beyond-FOV Large-Workspace Humanoid Manipulation with Omnidirectional 3D Perception
Pei Qu, Zheng Li, Yufei Jia +5
The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace.…
Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
Guoyang Zhao, Yudong Li, Weiqing Qi +5
Conventional navigation pipelines for legged robots remain largely geometry-centric, relying on dense SLAM representations that are fragile under rapid motion and offer limited sup…