Learning to Execute
arXiv:1410.4615
Abstract
Recurrent Neural Networks (RNNs) with Long Short-Term Memory units (LSTM) are widely used because they are expressive and are easy to train. Our interest lies in empirically evaluating the expressiveness and the learnability of LSTMs in the sequence-to-sequence regime by training them to evaluate short computer programs, a domain that has traditionally been seen as too complex for neural networks. We consider a simple class of programs that can be evaluated with a single left-to-right pass using constant memory. Our main result is that LSTMs can learn to map the character-level representations of such programs to their correct outputs. Notably, it was necessary to use curriculum learning, and while conventional curriculum learning proved ineffective, we developed a new variant of curriculum learning that improved our networks' performance in all experimental conditions. The improved curriculum had a dramatic impact on an addition problem, making it possible to train an LSTM to add two 9-digit numbers with 99% accuracy.
References in corpus (3)
Cited by in corpus (133)
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- A Critical Review of Recurrent Neural Networks for Sequence Learning
- Evaluating Large Language Models Trained on Code
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Text Understanding from Scratch
- Gated Feedback Recurrent Neural Networks
- Grammar as a Foreign Language
- Hindsight Experience Replay
- Grid Long Short-Term Memory
- Convolutional Neural Networks over Tree Structures for Programming Language Processing
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Beyond Short Snippets: Deep Networks for Video Classification
- Doctor AI: Predicting Clinical Events via Recurrent Neural Networks
- Learning Natural Language Inference using Bidirectional LSTM model and Inner-Attention
- Neural Programmer-Interpreters
- Neural GPUs Learn Algorithms
- End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures
- Sequence to Sequence -- Video to Text
- Long Short-Term Memory-Networks for Machine Reading
- Order Matters: Sequence to sequence for sets
- Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
- Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning
- Graph2Seq: Graph to Sequence Learning with Attention-based Neural Networks
- Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
- Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks
- Automatic Goal Generation for Reinforcement Learning Agents
- Deep Kalman Filters
- A Convolutional Attention Network for Extreme Summarization of Source Code
- Reverse Curriculum Generation for Reinforcement Learning
- Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings
- Self Paced Deep Learning for Weakly Supervised Object Detection
- Query-Efficient Imitation Learning for End-to-End Autonomous Driving
- DeepMath - Deep Sequence Models for Premise Selection
- Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Path
- Reinforcement Learning Neural Turing Machines - Revised
- RobustFill: Neural Program Learning under Noisy I/O
- Protein Secondary Structure Prediction with Long Short Term Memory Networks
- Curriculum Learning for Speech Emotion Recognition from Crowdsourced Labels
- Super Mario as a String: Platformer Level Generation Via LSTMs
- Character-Level Question Answering with Attention
- Analysing Mathematical Reasoning Abilities of Neural Models
- Noisy Activation Functions
- Video Summarization with Long Short-term Memory
- Hierarchical Attention Network for Action Recognition in Videos
- Learning Memory Access Patterns
- Improving performance of recurrent neural network with relu nonlinearity
- Strong error analysis for stochastic gradient descent optimization algorithms
- Learning to Navigate in Cities Without a Map
- The StreetLearn Environment and Dataset
- Deep Learning for Symbolic Mathematics
- ShapeWorld - A new test methodology for multimodal language understanding
- Discrete Flows: Invertible Generative Models of Discrete Data
- Automated Curriculum Learning for Neural Networks
- Deconfounding Reinforcement Learning in Observational Settings
- Automated curricula through setter-solver interactions
- What value do explicit high level concepts have in vision to language problems?
- Improving the Gating Mechanism of Recurrent Neural Networks
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Image Captioning with Deep Bidirectional LSTMs
- Program Synthesis with Large Language Models
- Mix&Match - Agent Curricula for Reinforcement Learning
- Neural Enquirer: Learning to Query Tables with Natural Language
- Relational recurrent neural networks
- Neural Program Synthesis with Priority Queue Training
- Energy-Based Hindsight Experience Prioritization
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- Attending to Mathematical Language with Transformers
- The Intentional Unintentional Agent: Learning to Solve Many Continuous Control Tasks Simultaneously
- SIGL: Securing Software Installations Through Deep Graph Learning
- Pointer Graph Networks
- Context-Aware Deep Spatio-Temporal Network for Hand Pose Estimation from Depth Images
- Generating Descriptions with Grounded and Co-Referenced People
- Attend to You: Personalized Image Captioning with Context Sequence Memory Networks
- Learning Typographic Style
- Visual Learning of Arithmetic Operations
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- Sequence to Sequence Learning for Optical Character Recognition
- Lie Access Neural Turing Machine
- Modeling Long-Range Context for Concurrent Dialogue Acts Recognition
- Extensions and Limitations of the Neural GPU
- State-Regularized Recurrent Neural Networks
- Deep Global-Relative Networks for End-to-End 6-DoF Visual Localization and Odometry
- A Survey on Curriculum Learning
- Measuring Arithmetic Extrapolation Performance
- Improved training for online end-to-end speech recognition systems
- Disentangled Representations in Neural Models
- Recognizing Implicit Discourse Relations via Repeated Reading: Neural Networks with Multi-Level Attention
- Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images
- Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval
- Towards Proof Synthesis Guided by Neural Machine Translation for Intuitionistic Propositional Logic
- When Do Curricula Work?
- Learning advanced mathematical computations from examples
- Learning to Sit: Synthesizing Human-Chair Interactions via Hierarchical Control
- Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
- Training Stronger Baselines for Learning to Optimize
- Coarse to Fine: Multi-label Image Classification with Global/Local Attention
- Article citation study: Context enhanced citation sentiment detection
- Curriculum By Smoothing
- Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent
- Hierarchical Recurrent Neural Network for Video Summarization
- Hierarchically Compositional Tasks and Deep Convolutional Networks
- Using Recurrent Neural Network for Learning Expressive Ontologies
- CFGs-2-NLU: Sequence-to-Sequence Learning for Mapping Utterances to Semantics and Pragmatics
- Real-time interactive sequence generation and control with Recurrent Neural Network ensembles
- MLBiNet: A Cross-Sentence Collective Event Detection Network
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- Interaction-limited Inverse Reinforcement Learning
- Neural Execution of Graph Algorithms
- Learning to Navigate the Web
- Towards Automatic Speech Identification from Vocal Tract Shape Dynamics in Real-time MRI
- Curriculum Design for Teaching via Demonstrations: Theory and Applications
- Learning and analyzing vector encoding of symbolic representations
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Lexicosyntactic Inference in Neural Models
- Learning Algorithmic Solutions to Symbolic Planning Tasks with a Neural Computer Architecture
- Survey of reasoning using Neural networks
- Region Growing Curriculum Generation for Reinforcement Learning
- Training a Subsampling Mechanism in Expectation
- An Empirical Comparison of Syllabuses for Curriculum Learning
- Online Visual Robot Tracking and Identification using Deep LSTM Networks
- Combining Learned Representations for Combinatorial Optimization
- On educating machines
- Progressive Growing of Neural ODEs
- Progress Extrapolating Algorithmic Learning to Arbitrary Sequence Lengths
- Meta-Learning an Inference Algorithm for Probabilistic Programs
- Latent Execution for Neural Program Synthesis
- Fixed -VAE Encoding for Curious Exploration in Complex 3D Environments
- Solving Graph-based Public Good Games with Tree Search and Imitation Learning
- Mastering Rate based Curriculum Learning
- Automated Curriculum Learning for Turn-level Spoken Language Understanding with Weak Supervision
- Growing Action Spaces
- Reinforced Temporal Attention and Split-Rate Transfer for Depth-Based Person Re-Identification