Open Scene Graphs for Open-World Object-Goal Navigation
arXiv:2508.04678 · doi:10.1177/02783649251369549
Abstract
How can we build general-purpose robot systems for open-world semantic navigation, e.g., searching a novel environment for a target object specified in natural language? To tackle this challenge, we introduce OSG Navigator, a modular system composed of foundation models, for open-world Object-Goal Navigation (ObjectNav). Foundation models provide enormous semantic knowledge about the world, but struggle to organise and maintain spatial information effectively at scale. Key to OSG Navigator is the Open Scene Graph representation, which acts as spatial memory for OSG Navigator. It organises spatial information hierarchically using OSG schemas, which are templates, each describing the common structure of a class of environments. OSG schemas can be automatically generated from simple semantic labels of a given environment, e.g., "home" or "supermarket". They enable OSG Navigator to adapt zero-shot to new environment types. We conducted experiments using both Fetch and Spot robots in simulation and in the real world, showing that OSG Navigator achieves state-of-the-art performance on ObjectNav benchmarks and generalises zero-shot over diverse goals, environments, and robot embodiments.
In IJRR Special Issue: Foundation Models and Neuro-symbolic AI for Robotics. Journal extension to arXiv:2407.02473
References in corpus (18)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- DINOv2: Learning Robust Visual Features without Supervision
- On Evaluation of Embodied Navigation Agents
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- Volumetric Instance-Aware Semantic Mapping and 3D Object Discovery
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
- Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
- L3MVN: Leveraging Large Language Models for Visual Target Navigation
- Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
- Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
- ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
- Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill
- Optimal Scene Graph Planning with Large Language Model Guidance
- Task and Motion Planning in Hierarchical 3D Scene Graphs
- Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation