On the Expressive Power of Deep Neural Networks
arXiv:1606.05336
Abstract
We propose a new approach to the problem of neural network expressivity, which seeks to characterize how structural properties of a neural network family affect the functions it is able to compute. Our approach is based on an interrelated set of measures of expressivity, unified by the novel notion of trajectory length, which measures how the output of a network changes as the input sweeps along a one-dimensional path. Our findings can be summarized as follows: (1) The complexity of the computed function grows exponentially with depth. (2) All weights are not equal: trained networks are more sensitive to their lower (initial) layer weights. (3) Regularizing on trajectory length (trajectory regularization) is a simpler alternative to batch normalization, with the same performance.
Accepted to ICML 2017
References in corpus (4)
Cited by in corpus (37)
- A trans-disciplinary review of deep learning research for water resources scientists
- Machine Learning and Deep Learning Algorithms for Bearing Fault Diagnostics -- A Comprehensive Review
- Searching for Activation Functions
- Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification
- A Closer Look at Memorization in Deep Networks
- ReLU Networks as Surrogate Models in Mixed-Integer Linear Programs
- Recovery Guarantees for One-hidden-layer Neural Networks
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Understanding Deep Neural Networks with Rectified Linear Units
- Hidden Fluid Mechanics: A Navier-Stokes Informed Deep Learning Framework for Assimilating Flow Visualization Data
- Deep Learning for Computational Chemistry
- Learning Non-overlapping Convolutional Neural Networks with Multiple Kernels
- On the Number of Linear Regions of Convolutional Neural Networks
- Adversarial Reprogramming of Neural Networks
- Technical Considerations for Semantic Segmentation in MRI using Convolutional Neural Networks
- Is Deeper Better only when Shallow is Good?
- Stein Neural Sampler
- The Upper Bound on Knots in Neural Networks
- Deep Function Machines: Generalized Neural Networks for Topological Layer Expression
- Improving GAN Training via Binarized Representation Entropy (BRE) Regularization
- Maximum Entropy Flow Networks
- End-to-end Learning of a Convolutional Neural Network via Deep Tensor Decomposition
- Gradient-based Training of Slow Feature Analysis by Differentiable Approximate Whitening
- The empirical size of trained neural networks
- Capacity allocation analysis of neural networks: A tool for principled architecture design
- Capacity allocation through neural network layers
- Investigating the Compositional Structure Of Deep Neural Networks
- On the Compressive Power of Deep Rectifier Networks for High Resolution Representation of Class Boundaries
- Deep Learning Works in Practice. But Does it Work in Theory?
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- Learning Boolean Circuits with Neural Networks
- Generalization Performance of Empirical Risk Minimization on Over-parameterized Deep ReLU Nets
- Revisiting hard thresholding for DNN pruning
- A lattice-based approach to the expressivity of deep ReLU neural networks
- Disentangled Neural Architecture Search
- Understanding the Importance of Single Directions via Representative Substitution
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal Matrices