Multiplicative LSTM for sequence modelling
arXiv:1609.07959
Abstract
We introduce multiplicative LSTM (mLSTM), a recurrent neural network architecture for sequence modelling that combines the long short-term memory (LSTM) and multiplicative recurrent neural network architectures. mLSTM is characterised by its ability to have different recurrent transition functions for each possible input, which we argue makes it more expressive for autoregressive density estimation. We demonstrate empirically that mLSTM outperforms standard LSTM and its deep variants for a range of character level language modelling tasks. In this version of the paper, we regularise mLSTM to achieve 1.27 bits/char on text8 and 1.24 bits/char on Hutter Prize. We also apply a purely byte-level mLSTM on the WikiText-2 dataset to achieve a character level entropy of 1.26 bits/char, corresponding to a word level perplexity of 88.8, which is comparable to word level LSTMs regularised in similar ways on the same task.
References in corpus (11)
- Pointer Sentinel Mixture Models
- Regularizing and Optimizing LSTM Language Models
- Learning to Generate Reviews and Discovering Sentiment
- Neural Machine Translation in Linear Time
- Hierarchical Multiscale Recurrent Neural Networks
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Improving Neural Language Models with a Continuous Cache
- On Multiplicative Integration with Recurrent Neural Networks
- Surprisal-Driven Feedback in Recurrent Networks
- Recurrent Memory Array Structures
Cited by in corpus (35)
- Learning to Generate Reviews and Discovering Sentiment
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- MeanSum: A Neural Model for Unsupervised Multi-document Abstractive Summarization
- Evaluating Protein Transfer Learning with TAPE
- Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
- Encoding word order in complex embeddings
- Dynamic Evaluation of Neural Sequence Models
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Deep Learning-based Sentiment Classification: A Comparative Survey
- Compressive Transformers for Long-Range Sequence Modelling
- Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL
- Practical Text Classification With Large Pre-Trained Language Models
- Dynamic Evaluation of Transformer Language Models
- Recurrent Additive Networks
- Who Needs Words? Lexicon-Free Speech Recognition
- Pre-Training of Deep Bidirectional Protein Sequence Representations with Structural Information
- Evaluating Text GANs as Language Models
- Neural Language Generation: Formulation, Methods, and Evaluation
- Combining Convolution and Recursive Neural Networks for Sentiment Analysis
- Deep Learning in Protein Structural Modeling and Design
- A Scalable Framework for Multilevel Streaming Data Analytics using Deep Learning
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Generating Sentiment-Preserving Fake Online Reviews Using Neural Language Models and Their Human- and Machine-based Detection
- Spatial PixelCNN: Generating Images from Patches
- Compressing Gradient Optimizers via Count-Sketches
- A Neural Virtual Anchor Synthesizer based on Seq2Seq and GAN Models
- Findings of the Second Workshop on Neural Machine Translation and Generation
- Multi-Zone Unit for Recurrent Neural Networks
- Melody Classification based on Performance Event Vector and BRNN
- Sparse Meta Networks for Sequential Adaptation and its Application to Adaptive Language Modelling
- Music Classification in MIDI Format based on LSTM Mdel
- Recognizing Long Grammatical Sequences Using Recurrent Networks Augmented With An External Differentiable Stack
- Modeling Multivariate Cyber Risks: Deep Learning Dating Extreme Value Theory
- Poly-NL: Linear Complexity Non-local Layers with Polynomials
- Long Short Term Memory Networks for Bandwidth Forecasting in Mobile Broadband Networks under Mobility