Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
arXiv:2203.12667 · doi:10.18653/v1/2022.acl-long.524
Abstract
A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks. Vision-and-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities. In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc. Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work. This paper serves as a thorough reference for the VLN research community.
19 pages. Accepted to ACL 2022
References in corpus (16)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Replica Dataset: A Digital Replica of Indoor Spaces
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
- From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following
- DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
- The StreetLearn Environment and Dataset
- Learning to Map Natural Language Instructions to Physical Quadcopter Control using Simulated Flight
- MeetUp! A Corpus of Joint Activity Dialogues in a Visual Environment
- RMM: A Recursive Mental Model for Dialog Navigation
- Deep Learning for Embodied Vision Navigation: A Survey
- The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation
- Diagnosing Vision-and-Language Navigation: What Really Matters
- Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation
- A Framework for Learning to Request Rich and Contextually Useful Information from Humans
- Rethinking the Spatial Route Prior in Vision-and-Language Navigation
Cited by in corpus (10)
- VLP: A Survey on Vision-Language Pre-training
- Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
- DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
- Open Scene Graphs for Open-World Object-Goal Navigation
- Language to Map: Topological map generation from natural language path instructions
- Hierarchical Path-planning from Speech Instructions with Spatial Concept-based Topometric Semantic Mapping
- Multimodal Speech Recognition for Language-Guided Embodied Agents
- Memory-Maze: Scenario Driven Visual Language Navigation Benchmark for Guiding Blind People
- Real-world Instance-specific Image Goal Navigation: Bridging Domain Gaps via Contrastive Learning
- Revolutionizing Turn-by-Turn Navigation with Cloud-Edge Deep Learning