Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
arXiv:1506.04089
Abstract
We propose a neural sequence-to-sequence model for direction following, a task that is essential to realizing effective autonomous agents. Our alignment-based encoder-decoder model with long short-term memory recurrent neural networks (LSTM-RNN) translates natural language instructions to action sequences based upon a representation of the observable world state. We introduce a multi-level aligner that empowers our model to focus on sentence "regions" salient to the current world state by using multiple abstractions of the input sentence. In contrast to existing methods, our model uses no specialized linguistic resources (e.g., parsers) or task-specific annotations (e.g., seed lexicons). It is therefore generalizable, yet still achieves the best results reported to-date on a benchmark single-sentence dataset and competitive results for the limited-training multi-sentence setting. We analyze our model through a series of ablations that elucidate the contributions of the primary components of our model.
To appear at AAAI 2016 (and an extended version of a NIPS 2015 Multimodal Machine Learning workshop paper)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Recurrent Neural Network Regularization
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Cited by in corpus (38)
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- Language to Logical Form with Neural Attention
- A Syntactic Neural Model for General-Purpose Code Generation
- A Benchmark for Systematic Generalization in Grounded Language Understanding
- Hierarchical Decision Making by Generating and Following Natural Language Instructions
- GPT3-to-plan: Extracting plans from text using GPT-3
- Data Recombination for Neural Semantic Parsing
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
- Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
- Look, Listen, and Act: Towards Audio-Visual Embodied Navigation
- Unified Pragmatic Models for Generating and Following Instructions
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning
- Mastering emergent language: learning to guide in simulated navigation
- Active Visual Information Gathering for Vision-Language Navigation
- Mapping Natural Language Instructions to Mobile UI Action Sequences
- Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation
- A Narration-based Reward Shaping Approach using Grounded Natural Language Commands
- ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
- Interactive Learning from Activity Description
- Pre-trained Word Embeddings for Goal-conditional Transfer Learning in Reinforcement Learning
- Executing Instructions in Situated Collaborative Interactions
- Language-guided Semantic Mapping and Mobile Manipulation in Partially Observable Environments
- SIGVerse: A cloud-based VR platform for research on social and embodied human-robot interaction
- Pre-Learning Environment Representations for Data-Efficient Neural Instruction Following
- Learning to Request Guidance in Emergent Communication
- Compositional Networks Enable Systematic Generalization for Grounded Language Understanding
- Towards Navigation by Reasoning over Spatial Configurations
- Grounding Complex Navigational Instructions Using Scene Graphs
- Building Intelligent Autonomous Navigation Agents
- Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question Answering
- Guided Policy Search for Parameterized Skills using Adverbs
- Machine Learning for Temporal Data in Finance: Challenges and Opportunities
- Task-Oriented Language Grounding for Language Input with Multiple Sub-Goals of Non-Linear Order
- From Route Instructions to Landmark Graphs
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout