Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
arXiv:2403.17846 · doi:10.15607/RSS.2024.XX.077
Abstract
Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features. While these maps allow for the prediction of point-wise saliency maps when queried for a certain language concept, large-scale environments and abstract queries beyond the object level still pose a considerable hurdle, ultimately limiting language-grounded robotic navigation. In this work, we present HOV-SG, a hierarchical open-vocabulary 3D scene graph mapping approach for language-grounded robot navigation. Leveraging open-vocabulary vision foundation models, we first obtain state-of-the-art open-vocabulary segment-level maps in 3D and subsequently construct a 3D scene graph hierarchy consisting of floor, room, and object concepts, each enriched with open-vocabulary features. Our approach is able to represent multi-story buildings and allows robotic traversal of those using a cross-floor Voronoi graph. HOV-SG is evaluated on three distinct datasets and surpasses previous baselines in open-vocabulary semantic accuracy on the object, room, and floor level while producing a 75% reduction in representation size compared to dense open-vocabulary maps. In order to prove the efficacy and generalization capabilities of HOV-SG, we showcase successful long-horizon language-conditioned robot navigation within real-world multi-storage environments. We provide code and trial video data at http://hovsg.github.io/.
Code and video are available at http://hovsg.github.io/
Cited by in corpus (17)
- Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
- Generative AI Agents in Autonomous Machines: A Safety Perspective
- Open Scene Graphs for Open-World Object-Goal Navigation
- DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes
- DynamicGSG: Dynamic 3D Gaussian Scene Graphs for Environment Adaptation
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
- REACT: Real-time Efficient Attribute Clustering and Transfer for Updatable 3D Scene Graph
- Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models
- SEMNAV: Enhancing Visual Semantic Navigation in Robotics through Semantic Segmentation
- Uncertainty-Informed Active Perception for Open Vocabulary Object Goal Navigation
- OSMa-Bench: Evaluating Open Semantic Mapping Under Varying Lighting Conditions
- Semantic Enrichment of CAD-Based Industrial Environments via Scene Graphs for Simulation and Reasoning
- Event-Grounding Graph: Unified Spatio-Temporal Scene Graph from Robotic Observations
- LAMP: Implicit Language Map for Robot Navigation
- Concept-Guided Exploration: Building Persistent, Actionable Scene Graphs
- Guessing human intentions to avoid dangerous situations in caregiving robots
- Do Visual-Language Grid Maps Capture Latent Semantics?