On the number of response regions of deep feed forward networks with piece-wise linear activations
arXiv:1312.6098
Abstract
This paper explores the complexity of deep feedforward networks with linear pre-synaptic couplings and rectified linear activations. This is a contribution to the growing body of work contrasting the representational power of deep and shallow network architectures. In particular, we offer a framework for comparing deep and shallow models that belong to the family of piecewise linear functions based on computational geometry. We look at a deep rectifier multi-layer perceptron (MLP) with linear outputs units and compare it with a single layer version of the model. In the asymptotic regime, when the number of inputs stays constant, if the shallow model has hidden units and inputs, then the number of linear regions is . For a layer model with hidden units on each layer it is . The number grows faster than when tends to infinity or when tends to infinity and . Additionally, even when is small, if we restrict to be , we can show that a deep model has considerably more linear regions that a shallow one. We consider this as a first step towards understanding the complexity of these models and specifically towards providing suitable mathematical tools for future analysis.
17 pages, 9 figures
References in corpus (4)
Cited by in corpus (60)
- On the Number of Linear Regions of Deep Neural Networks
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- How to Construct Deep Recurrent Neural Networks
- Understanding Deep Neural Networks with Rectified Linear Units
- Multifidelity Modeling for Physics-Informed Neural Networks (PINNs)
- On the Expressive Power of Deep Neural Networks
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- On Characterizing the Capacity of Neural Networks using Algebraic Topology
- Bounding and Counting Linear Regions of Deep Neural Networks
- Zero-bias autoencoders and the benefits of co-adapting features
- On the Number of Linear Regions of Convolutional Neural Networks
- A Tropical Approach to Neural Networks with Piecewise Linear Activations
- Is Deeper Better only when Shallow is Good?
- Some Theoretical Insights into Wasserstein GANs
- Discovering and Explaining the Representation Bottleneck of DNNs
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- Learning Lyapunov Functions for Piecewise Affine Systems with Neural Network Controllers
- Understanding Locally Competitive Networks
- Three-dimensional convolutional neural networks for neutrinoless double-beta decay signal/background discrimination in high-pressure gaseous Time Projection Chamber
- Linear predictor on linearly-generated data with missing values: non consistency and solutions
- Dissecting Deep Neural Networks
- Survey of Expressivity in Deep Neural Networks
- The Upper Bound on Knots in Neural Networks
- Training Neural Networks by Using Power Linear Units (PoLUs)
- Deep Narrow Boltzmann Machines are Universal Approximators
- An Embedding of ReLU Networks and an Analysis of their Identifiability
- Layer Folding: Neural Network Depth Reduction using Activation Linearization
- Provably Correct Training of Neural Network Controllers Using Reachability Analysis
- Sharp bounds for the number of regions of maxout networks and vertices of Minkowski sums
- Measuring Model Complexity of Neural Networks with Curve Activation Functions
- The empirical size of trained neural networks
- Feature selection of neural networks is skewed towards the less abstract cue
- Memory Matters: Convolutional Recurrent Neural Network for Scene Text Recognition
- Depth separation beyond radial functions
- Truncating Wide Networks using Binary Tree Architectures
- When Hardness of Approximation Meets Hardness of Learning
- On the Number of Linear Functions Composing Deep Neural Network: Towards a Refined Definition of Neural Networks Complexity
- From Deep to Shallow: Transformations of Deep Rectifier Networks
- Rethink the Connections among Generalization, Memorization and the Spectral Bias of DNNs
- Investigating the Compositional Structure Of Deep Neural Networks
- A Tale of Three Probabilistic Families: Discriminative, Descriptive and Generative Models
- In Proximity of ReLU DNN, PWA Function, and Explicit MPC
- Knots in random neural networks
- Neural networks behave as hash encoders: An empirical study
- Learning Boolean Circuits with Neural Networks
- Deep Reinforcement Learning with Linear Quadratic Regulator Regions
- Improving learnability of neural networks: adding supplementary axes to disentangle data representation
- A simple geometric proof for the benefit of depth in ReLU networks
- Fast Jacobian-Vector Product for Deep Networks
- How to Explain Neural Networks: an Approximation Perspective
- Learning Lyapunov Functions for Hybrid Systems
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- Light Multi-segment Activation for Model Compression
- Interpretations of Deep Learning by Forests and Haar Wavelets
- Unsupervised Representation Learning via Neural Activation Coding
- Tucker Decomposition Network: Expressive Power and Comparison
- Iterative evaluation of LSTM cells
- How Could Polyhedral Theory Harness Deep Learning?
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- The emergent algebraic structure of RNNs and embeddings in NLP