Learning the Number of Neurons in Deep Networks
arXiv:1611.06321
Abstract
Nowadays, the number of layers and of neurons in each layer of a deep network are typically set manually. While very deep and wide networks have proven effective in general, they come at a high memory and computation cost, thus making them impractical for constrained platforms. These networks, however, are known to have many redundant parameters, and could thus, in principle, be replaced by more compact architectures. In this paper, we introduce an approach to automatically determining the number of neurons in each layer of a deep network during learning. To this end, we propose to make use of structured sparsity during learning. More precisely, we use a group sparsity regularizer on the parameters of the network, where each group is defined to act on a single neuron. Starting from an overcomplete network, we show that our approach can reduce the number of parameters by up to 80\% while retaining or even improving the network accuracy.
NIPS 2016
Cited by in corpus (71)
- Neural Granger Causality
- Deep Learning for Generic Object Detection: A Survey
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Network Pruning via Transformable Architecture Search
- MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning
- Deep Learning Methods for Solving Linear Inverse Problems: Research Directions and Paradigms
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- Distilling Object Detectors via Decoupled Features
- Computationally Efficient Neural Image Compression
- Structural Compression of Convolutional Neural Networks
- Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression
- Bringing Giant Neural Networks Down to Earth with Unlabeled Data
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
- DHP: Differentiable Meta Pruning via HyperNetworks
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Equivalent and Approximate Transformations of Deep Neural Networks
- PruneNet: Channel Pruning via Global Importance
- Depth-wise Decomposition for Accelerating Separable Convolutions in Efficient Convolutional Neural Networks
- Transformed Regularization for Learning Sparse Deep Neural Networks
- Learning Filter Basis for Convolutional Neural Network Compression
- Consistent Sparse Deep Learning: Theory and Computation
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Learning Strict Identity Mappings in Deep Residual Networks
- Filter Sketch for Network Pruning
- VACL: Variance-Aware Cross-Layer Regularization for Pruning Deep Residual Networks
- Accelerate CNN via Recursive Bayesian Pruning
- Generalized Bayesian Posterior Expectation Distillation for Deep Neural Networks
- Do We Need Fully Connected Output Layers in Convolutional Networks?
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- Separable Layers Enable Structured Efficient Linear Substitutions
- Refining the Structure of Neural Networks Using Matrix Conditioning
- Gradient-Coherent Strong Regularization for Deep Neural Networks
- HALP: Hardware-Aware Latency Pruning
- When to Prune? A Policy towards Early Structural Pruning
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- ADMP: An Adversarial Double Masks Based Pruning Framework For Unsupervised Cross-Domain Compression
- Training Sparse Neural Networks using Compressed Sensing
- Filter Pruning using Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks
- Localization-aware Channel Pruning for Object Detection
- Improving Network Slimming with Nonconvex Regularization
- Sparsely Grouped Input Variables for Neural Networks
- Efficient Proximal Mapping of the 1-path-norm of Shallow Networks
- Neural Architecture Search as Sparse Supernet
- Green DetNet: Computation and Memory efficient DetNet using Smart Compression and Training
- BEAN: Interpretable Representation Learning with Biologically-Enhanced Artificial Neuronal Assembly Regularization
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- Deep Stacked Stochastic Configuration Networks for Lifelong Learning of Non-Stationary Data Streams
- Fuzzy Logic Interpretation of Quadratic Networks
- Distilling Image Classifiers in Object Detectors
- -LBI: Stochastic Split Linearized Bregman Iterations for Parsimonious Deep Learning
- An Improving Framework of regularization for Network Compression
- Regularization and Reparameterization Avoid Vanishing Gradients in Sigmoid-Type Networks
- Content-Aware Convolutional Neural Networks
- Channel Planting for Deep Neural Networks using Knowledge Distillation
- On tuning deep learning models: a data mining perspective
- Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks
- Out-of-the-box channel pruned networks
- Embedding Differentiable Sparsity into Deep Neural Network
- A Partial Regularization Method for Network Compression
- Self-grouping Convolutional Neural Networks
- Training CNNs faster with Dynamic Input and Kernel Downsampling
- DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths
- C2S2: Cost-aware Channel Sparse Selection for Progressive Network Pruning
- NodeDrop: A Condition for Reducing Network Size without Effect on Output
- Alternate Model Growth and Pruning for Efficient Training of Recommendation Systems
- Schematic Memory Persistence and Transience for Efficient and Robust Continual Learning