Adaptive Input Representations for Neural Language Modeling
arXiv:1809.10853
Abstract
We introduce adaptive input representations for neural language modeling which extend the adaptive softmax of Grave et al. (2017) to input representations of variable capacity. There are several choices on how to factorize the input and output layers, and whether to model words, characters or sub-word units. We perform a systematic comparison of popular choices for a self-attentional architecture. Our experiments show that models equipped with adaptive embeddings are more than twice as fast to train than the popular character input CNN while having a lower number of parameters. On the WikiText-103 benchmark we achieve 18.7 perplexity, an improvement of 10.5 perplexity compared to the previously best published result and on the Billion Word benchmark, we achieve 23.02 perplexity.
12 pages
References in corpus (5)
Cited by in corpus (21)
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Reducing Transformer Depth on Demand with Structured Dropout
- Training with Quantization Noise for Extreme Model Compression
- Sparse Sinkhorn Attention
- Compressive Transformers for Long-Range Sequence Modelling
- Tensorized Embedding Layers for Efficient Model Compression
- TNT-KID: Transformer-based Neural Tagger for Keyword Identification
- Language Models with Transformers
- Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs
- One Epoch Is All You Need
- Primer: Searching for Efficient Transformers for Language Modeling
- OmniNet: Omnidirectional Representations from Transformers
- Parameter-Efficient Transfer from Sequential Behaviors for User Modeling and Recommendation
- A Generic Network Compression Framework for Sequential Recommender Systems
- PopMAG: Pop Music Accompaniment Generation
- Extractive Summary as Discrete Latent Variables
- Stable Invariant Models via Koopman Spectra
- Capturing Structural Locality in Non-parametric Language Models
- Normalization of Input-output Shared Embeddings in Text Generation Models