Conditional Computation in Neural Networks for faster models
arXiv:1511.06297
Abstract
Deep learning has become the state-of-art tool in many applications, but the evaluation and training of deep models can be time-consuming and computationally expensive. The conditional computation approach has been proposed to tackle this problem (Bengio et al., 2013; Davis & Arel, 2013). It operates by selectively activating only parts of the network at a time. In this paper, we use reinforcement learning as a tool to optimize conditional computation policies. More specifically, we cast the problem of learning activation-dependent policies for dropping out blocks of units as a reinforcement learning problem. We propose a learning scheme motivated by computation speed, capturing the idea of wanting to have parsimonious activations while maintaining prediction accuracy. We apply a policy gradient algorithm for learning policies that optimize this loss function and propose a regularization mechanism that encourages diversification of the dropout policy. We present encouraging empirical results showing that this approach improves the speed of computation without impacting the quality of the approximation.
ICLR 2016 submission, revised
References in corpus (3)
Cited by in corpus (64)
- Adaptive Computation Time for Recurrent Neural Networks
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Learning Sparse Neural Networks through Regularization
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- Adaptive Neural Networks for Efficient Inference
- XNOR-Net++: Improved Binary Neural Networks
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- Learning model-based planning from scratch
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Learning to Continually Learn
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- MaskConnect: Connectivity Learning by Gradient Descent
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Dynamic Neural Networks: A Survey
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Simple, Distributed, and Accelerated Probabilistic Programming
- Routing Networks and the Challenges of Modular and Compositional Computation
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- Scaling Vision with Sparse Mixture of Experts
- Attention over Parameters for Dialogue Systems
- Spatially Adaptive Computation Time for Residual Networks
- Autoencoding with a Classifier System
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
- The Tree Ensemble Layer: Differentiability meets Conditional Computation
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
- Faster Asynchronous SGD
- Biased Mixtures Of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
- Dynamic Deep Neural Networks: Optimizing Accuracy-Efficiency Trade-offs by Selective Execution
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
- Controlling Computation versus Quality for Neural Sequence Models
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Deep Convolutional Decision Jungle for Image Classification
- Flexible Multi-task Networks by Learning Parameter Allocation
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- Conditional Computation for Continual Learning
- InstaNAS: Instance-aware Neural Architecture Search
- Surprisal-Triggered Conditional Computation with Neural Networks
- VA-RED: Video Adaptive Redundancy Reduction
- You Look Twice: GaterNet for Dynamic Filter Selection in CNNs
- Deep Online Convex Optimization with Gated Games
- Adaptive Memory Networks
- Energy-efficient Amortized Inference with Cascaded Deep Classifiers
- Learning Sparse Mixture of Experts for Visual Question Answering
- Dynamic Network Quantization for Efficient Video Inference
- Multi-modal Experts Network for Autonomous Driving
- TreeSegNet: Adaptive Tree CNNs for Subdecimeter Aerial Image Segmentation
- High-Capacity Expert Binary Networks
- Not All Attention Is Needed: Gated Attention Network for Sequence Data
- Question Guided Modular Routing Networks for Visual Question Answering
- Interpretable Neural Network Decoupling
- Shapley Explanation Networks
- URNet : User-Resizable Residual Networks with Conditional Gating Module
- Toward Runtime-Throttleable Neural Networks
- Deep Learning with a Classifier System: Initial Results
- Depth-Adaptive Graph Recurrent Network for Text Classification
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- Layer Flexible Adaptive Computational Time
- Unbiased Gradient Estimation with Balanced Assignments for Mixtures of Experts
- Robust Text Classifier on Test-Time Budgets
- Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference
- CNN with large memory layers
- Restricted Recurrent Neural Tensor Networks: Exploiting Word Frequency and Compositionality