Direct Output Connection for a High-Rank Language Model
arXiv:1808.10143
Abstract
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers. Our proposed method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. (2018). The proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets. Moreover, we indicate our proposed method contributes to two application tasks: machine translation and headline generation. Our code is publicly available at: https://github.com/nttcslab-nlp/doc_lm.
EMNLP 2018 paper
References in corpus (11)
- Sequence to Sequence Learning with Neural Networks
- Neural Architecture Search with Reinforcement Learning
- Recurrent Neural Network Regularization
- Pointer Sentinel Mixture Models
- Regularizing and Optimizing LSTM Language Models
- Selective Encoding for Abstractive Sentence Summarization
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Constituency Parsing with a Self-Attentive Encoder
- Source-side Prediction for Neural Headline Generation
- Improving Neural Parsing by Disentangling Model Combination and Reranking Effects
- Fraternal Dropout