Grid Long Short-Term Memory
arXiv:1507.01526
Abstract
This paper introduces Grid Long Short-Term Memory, a network of LSTM cells arranged in a multidimensional grid that can be applied to vectors, sequences or higher dimensional data such as images. The network differs from existing deep LSTM architectures in that the cells are connected between network layers as well as along the spatiotemporal dimensions of the data. The network provides a unified way of using LSTM for both deep and sequential computation. We apply the model to algorithmic tasks such as 15-digit integer addition and sequence memorization, where it is able to significantly outperform the standard LSTM. We then give results for two empirical tasks. We find that 2D Grid LSTM achieves 1.47 bits per character on the Wikipedia character prediction benchmark, which is state-of-the-art among neural approaches. In addition, we use the Grid LSTM to define a novel two-dimensional translation model, the Reencoder, and show that it outperforms a phrase-based reference system on a Chinese-to-English translation task.
15 pages
References in corpus (8)
- Adam: A Method for Stochastic Optimization
- Sequence to Sequence Learning with Neural Networks
- LSTM: A Search Space Odyssey
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Fractional Max-Pooling
- ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks
- Spatially-sparse convolutional neural networks
- Multi-column Deep Neural Networks for Image Classification
Cited by in corpus (78)
- Conditional Image Generation with PixelCNN Decoders
- Pixel Recurrent Neural Networks
- Training Very Deep Networks
- Resnet in Resnet: Generalizing Residual Architectures
- Recent Advances in Recurrent Neural Networks
- Prediction of Sea Surface Temperature using Long Short-Term Memory
- Adaptive Computation Time for Recurrent Neural Networks
- Hierarchical Multiscale Recurrent Neural Networks
- Neural GPUs Learn Algorithms
- Long Short-Term Memory-Networks for Machine Reading
- Convolutional Neural Fabrics
- Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
- Stabilizing Transformers for Reinforcement Learning
- MelNet: A Generative Model for Audio in the Frequency Domain
- Multiplicative LSTM for sequence modelling
- A Radar Signal Deinterleaving Method Based on Semantic Segmentation with Neural Network
- On Multiplicative Integration with Recurrent Neural Networks
- Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Prediction
- A Review of 40 Years of Cognitive Architecture Research: Core Cognitive Abilities and Practical Applications
- An End-to-End Breast Tumour Classification Model Using Context-Based Patch Modelling- A BiLSTM Approach for Image Classification
- STFCN: Spatio-Temporal FCN for Semantic Video Segmentation
- On the State of the Art of Evaluation in Neural Language Models
- Dual Rectified Linear Units (DReLUs): A Replacement for Tanh Activation Functions in Quasi-Recurrent Neural Networks
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks
- Visualizing and Understanding Curriculum Learning for Long Short-Term Memory Networks
- Semantic Object Parsing with Local-Global Long Short-Term Memory
- Highway Long Short-Term Memory RNNs for Distant Speech Recognition
- Learning Efficient Algorithms with Hierarchical Attentive Memory
- Learning Affinity via Spatial Propagation Networks
- Improving the Neural GPU Architecture for Algorithm Learning
- Recurrent Memory Networks for Language Modeling
- You May Not Need Attention
- Attending to Mathematical Language with Transformers
- Embedding API Dependency Graph for Neural Code Generation
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Recurrent Attention Unit
- Neural Shuffle-Exchange Networks -- Sequence Processing in O(n log n) Time
- Network Level Spatial Temporal Traffic State Forecasting with Hierarchical-Attention-LSTM (HierAttnLSTM)
- Deep Neural Machine Translation with Linear Associative Unit
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- Modelling Interaction of Sentence Pair with coupled-LSTMs
- Learning Generalizable Visual Representations via Interactive Gameplay
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Lie Access Neural Turing Machine
- Extensions and Limitations of the Neural GPU
- Geometric Scene Parsing with Hierarchical LSTM
- How deep learning works --The geometry of deep learning
- Learning Online Alignments with Continuous Rewards Policy Gradient
- A Distributed Neural Network Architecture for Robust Non-Linear Spatio-Temporal Prediction
- An Empirical Exploration of Skip Connections for Sequential Tagging
- Learning Deep Matrix Representations
- Face Parsing via Recurrent Propagation
- Neural Arithmetic Expression Calculator
- Semantic Object Parsing with Graph LSTM
- Surprisal-Driven Zoneout
- Low-rank passthrough neural networks
- A Latent Feelings-aware RNN Model for User Churn Prediction with Behavioral Data
- Distributed Sequence Memory of Multidimensional Inputs in Recurrent Networks
- Surprisal-Driven Feedback in Recurrent Networks
- Teaching Machines to Converse
- Recurrent Memory Array Structures
- Cell-aware Stacked LSTMs for Modeling Sentences
- Scene Labeling using Gated Recurrent Units with Explicit Long Range Conditioning
- ContextVP: Fully Context-Aware Video Prediction
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Short-term load forecasting using optimized LSTM networks based on EMD
- Context-Aware Sequence-to-Sequence Models for Conversational Systems
- Shortcut Sequence Tagging
- High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
- Graph2Kernel Grid-LSTM: A Multi-Cued Model for Pedestrian Trajectory Prediction by Learning Adaptive Neighborhoods
- Learning User Intent from Action Sequences on Interactive Systems
- Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource Languages
- High Order Recurrent Neural Networks for Acoustic Modelling
- A Novel Framework for Recurrent Neural Networks with Enhancing Information Processing and Transmission between Units
- MCRM: Mother Compact Recurrent Memory
- Deep Differential Recurrent Neural Networks
- Progressively Diffused Networks for Semantic Image Segmentation
- ARMA Nets: Expanding Receptive Field for Dense Prediction