Dynamic Capacity Networks
arXiv:1511.07838
Abstract
We introduce the Dynamic Capacity Network (DCN), a neural network that can adaptively assign its capacity across different portions of the input data. This is achieved by combining modules of two types: low-capacity sub-networks and high-capacity sub-networks. The low-capacity sub-networks are applied across most of the input, but also provide a guide to select a few portions of the input on which to apply the high-capacity sub-networks. The selection is made using a novel gradient-based attention mechanism, that efficiently identifies input regions for which the DCN's output is most sensitive and to which we should devote more capacity. We focus our empirical evaluation on the Cluttered MNIST and SVHN image datasets. Our findings indicate that DCNs are able to drastically reduce the number of computations, compared to traditional convolutional neural networks, while maintaining similar or even better performance.
ICML 2016
References in corpus (17)
- Adam: A Method for Stochastic Optimization
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
- Going Deeper with Convolutions
- Theano: new features and speed improvements
- Compressing Deep Convolutional Networks using Vector Quantization
- Recurrent Models of Visual Attention
- DRAW: A Recurrent Neural Network For Image Generation
- Multiple Object Recognition with Visual Attention
- Multi-digit Number Recognition from Street View Imagery using Deep Convolutional Neural Networks
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Blocks and Fuel: Frameworks for deep learning
- Deep Learning of Representations: Looking Forward
- Spatial Transformer Networks
- On Learning Where To Look
Cited by in corpus (6)
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Attention-guided Chained Context Aggregation for Semantic Segmentation
- Spatially Adaptive Computation Time for Residual Networks
- Person Re-identification Using Visual Attention
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Attend To Count: Crowd Counting with Adaptive Capacity Multi-scale CNNs