A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution
arXiv:2107.05612
Abstract
Natural language provides an accessible and expressive interface to specify long-term tasks for robotic agents. However, non-experts are likely to specify such tasks with high-level instructions, which abstract over specific robot actions through several layers of abstraction. We propose that key to bridging this gap between language and robot actions over long execution horizons are persistent representations. We propose a persistent spatial semantic representation method, and show how it enables building an agent that performs hierarchical reasoning to effectively execute long-term tasks. We evaluate our approach on the ALFRED benchmark and achieve state-of-the-art results, despite completely avoiding the commonly used step-by-step instructions.
Presented at CoRL 2021
References in corpus (5)
- Grounded Language Learning in a Simulated 3D World
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- A Joint Model of Language and Perception for Grounded Attribute Learning
- Learning to Map Natural Language Instructions to Physical Quadcopter Control using Simulated Flight
- Learning Models for Following Natural Language Directions in Unknown Environments
Cited by in corpus (5)
- CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
- FILM: Following Instructions in Language with Modular Methods
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- LUMINOUS: Indoor Scene Generation for Embodied AI Challenges
- Are you doing what I say? On modalities alignment in ALFRED