Efficient softmax approximation for GPUs
arXiv:1609.04309
Abstract
We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word distribution to form clusters that explicitly minimize the expectation of computation time. Our approach further reduces the computational time by exploiting the specificities of modern architectures and matrix-matrix vector operations, making it particularly suited for graphical processing units. Our experiments carried out on standard benchmarks, such as EuroParl and One Billion Word, show that our approach brings a large gain in efficiency over standard approximations while achieving an accuracy close to that of the full softmax. The code of our method is available at https://github.com/facebookresearch/adaptive-softmax.
Accepted to ICML 2017
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Exploring the Limits of Language Modeling
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Learning Longer Memory in Recurrent Neural Networks
- BlackOut: Speeding up Recurrent Neural Network Language Models With Very Large Vocabularies
- Strategies for Training Large Vocabulary Neural Language Models
- Learning Visual Features from Large Weakly Supervised Data
Cited by in corpus (22)
- HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Vocabulary Selection Strategies for Neural Machine Translation
- Automatic Speech Recognition with Very Large Conversational Finnish and Estonian Vocabularies
- Who Needs Words? Lexicon-Free Speech Recognition
- Fast Parametric Learning with Activation Memorization
- Adversarial Contrastive Estimation
- Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation
- Towards Understanding Neural Machine Translation with Word Importance
- Accelerated Training for Massive Classification via Dynamic Class Selection
- Lightweight Adaptive Mixture of Neural and N-gram Language Models
- Unsupervised and Efficient Vocabulary Expansion for Recurrent Neural Network Language Models in ASR
- Prototype Memory for Large-scale Face Representation Learning
- Augment and Reduce: Stochastic Inference for Large Categorical Distributions
- Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language Generation
- Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference
- Attention-based sequence-to-sequence model for speech recognition: development of state-of-the-art system on LibriSpeech and its application to non-native English
- Cavs: A Vertex-centric Programming Interface for Dynamic Neural Networks
- Improved training of neural trans-dimensional random field language models with dynamic noise-contrastive estimation
- The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles
- Integrating Discrete and Neural Features via Mixed-feature Trans-dimensional Random Field Language Models
- Error-Correcting Neural Sequence Prediction