Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
arXiv:2403.14163 · doi:10.1016/j.aei.2025.103135
Abstract
Object-goal navigation is a crucial engineering task for the community of embodied navigation; it involves navigating to an instance of a specified object category within unseen environments. Although extensive investigations have been conducted on both end-to-end and modular-based, data-driven approaches, fully enabling an agent to comprehend the environment through perceptual knowledge and perform object-goal navigation as efficiently as humans remains a significant challenge. Recently, large language models have shown potential in this task, thanks to their powerful capabilities for knowledge extraction and integration. In this study, we propose a data-driven, modular-based approach, trained on a dataset that incorporates common-sense knowledge of object-to-room relationships extracted from a large language model. We utilize the multi-channel Swin-Unet architecture to conduct multi-task learning incorporating with multimodal inputs. The results in the Habitat simulator demonstrate that our framework outperforms the baseline by an average of 10.6% in the efficiency metric, Success weighted by Path Length (SPL). The real-world demonstration shows that the proposed approach can efficiently conduct this task by traversing several rooms. For more details and real-world demonstrations, please check our project webpage (https://sunleyuan.github.io/ObjectNav).
will soon submit to the Elsevier journal, Advanced Engineering Informatics
References in corpus (19)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Graph Transformer Networks
- Learning to Explore using Active Neural SLAM
- Object Goal Navigation using Goal-Oriented Semantic Exploration
- L3MVN: Leveraging Large Language Models for Visual Target Navigation
- ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings
- Large Language Models for Robotics: A Survey
- An Online Semantic Mapping System for Extending and Enhancing Visual SLAM
- Frontier Semantic Exploration for Visual Target Navigation
- TransFusionOdom: Interpretable Transformer-based LiDAR-Inertial Fusion Odometry Estimation
- VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
- Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
- Traj-LIO: A Resilient Multi-LiDAR Multi-IMU State Estimator Through Sparse Gaussian Process
- Visual Semantic Navigation with Real Robots
- Advances in Embodied Navigation Using Large Language Models: A Survey
- Navigation with VLM framework: Towards Going to Any Language
- VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation