Surprisal-Triggered Conditional Computation with Neural Networks
arXiv:2006.01659
Abstract
Autoregressive neural network models have been used successfully for sequence generation, feature extraction, and hypothesis scoring. This paper presents yet another use for these models: allocating more computation to more difficult inputs. In our model, an autoregressive model is used both to extract features and to predict observations in a stream of input observations. The surprisal of the input, measured as the negative log-likelihood of the current observation according to the autoregressive model, is used as a measure of input difficulty. This in turn determines whether a small, fast network, or a big, slow network, is used. Experiments on two speech recognition tasks show that our model can match the performance of a baseline in which the big network is always used with 15% fewer FLOPs.
References in corpus (14)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Language Models are Few-Shot Learners
- Recurrent Neural Network Regularization
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- The Cost of Training NLP Models: A Concise Overview
- A Cascade Architecture for Keyword Spotting on Mobile Devices
- Variable Computation in Recurrent Neural Networks
- Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
- Changing Model Behavior at Test-Time Using Reinforcement Learning
- Surprisal-Driven Zoneout
- The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information Budget
- Surprisal-Driven Feedback in Recurrent Networks
- Flexible Deep Neural Network Processing