Pointer Sentinel Mixture Models
arXiv:1609.07843
Abstract
Recent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies. Even then they struggle to predict rare or unseen words even if the context makes the prediction unambiguous. We introduce the pointer sentinel mixture architecture for neural sequence models which has the ability to either reproduce a word from the recent context or produce a word from a standard softmax classifier. Our pointer sentinel-LSTM model achieves state of the art language modeling performance on the Penn Treebank (70.9 perplexity) while using far fewer parameters than a standard softmax LSTM. In order to evaluate how well language models can exploit longer contexts and deal with more realistic vocabularies and larger corpora we also introduce the freely available WikiText corpus.
References in corpus (2)
Cited by in corpus (71)
- Neural Architecture Search with Reinforcement Learning
- Regularizing and Optimizing LSTM Language Models
- Quasi-Recurrent Neural Networks
- Reducing Transformer Depth on Demand with Structured Dropout
- Neural Text Generation with Unlikelihood Training
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Data Noising as Smoothing in Neural Network Language Models
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- A Neural Knowledge Language Model
- Dynamic Evaluation of Neural Sequence Models
- Generalization through Memorization: Nearest Neighbor Language Models
- Challenges in Data-to-Document Generation
- Compressive Transformers for Long-Range Sequence Modelling
- On the State of the Art of Evaluation in Neural Language Models
- Low Resource Text Classification with ULMFit and Backtranslation
- Revisiting Activation Regularization for Language RNNs
- Cortical microcircuits as gated-recurrent neural networks
- Machine Reading Comprehension: a Literature Review
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Modeling Vocabulary for Big Code Machine Learning
- Attention-based Modeling for Emotion Detection and Classification in Textual Conversations
- Adaptively Truncating Backpropagation Through Time to Control Gradient Bias
- Implicit Dimension Identification in User-Generated Text with LSTM Networks
- On Generalization Bounds of a Family of Recurrent Neural Networks
- Can I trust you more? Model-Agnostic Hierarchical Explanations
- Domain-specific Communication Optimization for Distributed DNN Training
- A Neural Language Model for Dynamically Representing the Meanings of Unknown Words and Entities in a Discourse
- Unbounded cache model for online language modeling with open vocabulary
- Differentiable Architecture Search with Ensemble Gumbel-Softmax
- A Flexible Approach to Automated RNN Architecture Generation
- AMR Parsing as Sequence-to-Graph Transduction
- Knowledge-Augmented Language Model and its Application to Unsupervised Named-Entity Recognition
- Statistical Adaptive Stochastic Gradient Methods
- A Convergence Analysis for A Class of Practical Variance-Reduction Stochastic Gradient MCMC
- Learning to Attend, Copy, and Generate for Session-Based Query Suggestion
- Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
- Sparseout: Controlling Sparsity in Deep Networks
- FTRANS: Energy-Efficient Acceleration of Transformers using FPGA
- Compressing Deep Neural Networks via Layer Fusion
- Compressing Gradient Optimizers via Count-Sketches
- AdaSGD: Bridging the gap between SGD and Adam
- Be Concise and Precise: Synthesizing Open-Domain Entity Descriptions from Facts
- Pun Generation with Surprise
- Early Improving Recurrent Elastic Highway Network
- Transformer on a Diet
- FineText: Text Classification via Attention-based Language Model Fine-tuning
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- Context Models for OOV Word Translation in Low-Resource Languages
- Microsoft AI Challenge India 2018: Learning to Rank Passages for Web Question Answering with Deep Attention Networks
- Word Embedding based on Low-Rank Doubly Stochastic Matrix Decomposition
- Variational Smoothing in Recurrent Neural Network Language Models
- Learning Less-Overlapping Representations
- Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention
- Discovering Useful Sentence Representations from Large Pretrained Language Models
- Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference
- Character n-gram Embeddings to Improve RNN Language Models
- Learning to Adaptively Scale Recurrent Neural Networks
- Do Transformers Need Deep Long-Range Memory
- A Lightweight Recurrent Network for Sequence Modeling
- Distributional Discrepancy: A Metric for Unconditional Text Generation
- Sparse Meta Networks for Sequential Adaptation and its Application to Adaptive Language Modelling
- Topic, Sentiment and Impact Analysis: COVID19 Information Seeking on Social Media
- Where's My Head? Definition, Dataset and Models for Numeric Fused-Heads Identification and Resolution
- SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling
- Classification as Decoder: Trading Flexibility for Control in Medical Dialogue
- Input-to-Output Gate to Improve RNN Language Models
- Exchangeable modelling of relational data: checking sparsity, train-test splitting, and sparse exchangeable Poisson matrix factorization
- OrderNet: Ordering by Example
- Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
- Low-Complexity LSTM Training and Inference with FloatSD8 Weight Representation