Strategies for Training Large Vocabulary Neural Language Models
arXiv:1512.04906
Abstract
Training neural network language models over large vocabularies is still computationally very costly compared to count-based models such as Kneser-Ney. At the same time, neural language models are gaining popularity for many applications such as speech recognition and machine translation whose success depends on scalability. We present a systematic comparison of strategies to represent and train large vocabularies, including softmax, hierarchical softmax, target sampling, noise contrastive estimation and self normalization. We further extend self normalization to be a proper estimator of likelihood and introduce an efficient variant of softmax. We evaluate each method on three popular benchmarks, examining performance on rare words, the speed/accuracy trade-off and complementarity to Kneser-Ney.
12 pages; journal paper; under review
References in corpus (3)
Cited by in corpus (16)
- Exploring Sparsity in Recurrent Neural Networks
- Efficient softmax approximation for GPUs
- Recurrent Memory Networks for Language Modeling
- Fast Parametric Learning with Activation Memorization
- Compressing Neural Language Models by Sparse Word Representations
- Language Detection Engine for Multilingual Texting on Mobile Devices
- Accelerated Training for Massive Classification via Dynamic Class Selection
- The Z-loss: a shift and scale invariant classification loss belonging to the Spherical Family
- Softmax Dissection: Towards Understanding Intra- and Inter-class Objective for Embedding Learning
- Accelerating Large Scale Knowledge Distillation via Dynamic Importance Sampling
- Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference
- Syllable-level Neural Language Model for Agglutinative Language
- A Survey on Green Deep Learning
- Error-Correcting Neural Sequence Prediction
- Improving Word Representations: A Sub-sampled Unigram Distribution for Negative Sampling
- Restricted Recurrent Neural Tensor Networks: Exploiting Word Frequency and Compositionality