Representation Benefits of Deep Feedforward Networks
arXiv:1509.08101
Abstract
This note provides a family of classification problems, indexed by a positive integer , where all shallow networks with fewer than exponentially (in ) many nodes exhibit error at least , whereas a deep network with 2 nodes in each of layers achieves zero error, as does a recurrent network with 3 distinct nodes iterated times. The proof is elementary, and the networks are standard feedforward networks with ReLU (Rectified Linear Unit) nonlinearities.
References in corpus (1)
Cited by in corpus (66)
- Exponential expressivity in deep neural networks through transient chaos
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- The Power of Depth for Feedforward Neural Networks
- A Review on Deep Learning in Medical Image Reconstruction
- The Modern Mathematics of Deep Learning
- On the Expressive Power of Deep Learning: A Tensor Analysis
- Approximating Continuous Functions by ReLU Nets of Minimal Width
- Understanding Deep Neural Networks with Rectified Linear Units
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- Learning Functions: When Is Deep Better Than Shallow
- Nonlinear Approximation and (Deep) ReLU Networks
- DNN Expression Rate Analysis of High-dimensional PDEs: Application to Option Pricing
- Benefits of depth in neural networks
- Multifidelity Modeling for Physics-Informed Neural Networks (PINNs)
- On the Expressive Power of Deep Neural Networks
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Inductive Bias of Deep Convolutional Networks through Pooling Geometry
- Expressiveness of Rectifier Networks
- Artificial Intelligence in Materials Science and Engineering: Current Landscape, Key Challenges, and Future Trajectorie
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- Automated Architecture Design for Deep Neural Networks
- On the Number of Linear Regions of Convolutional Neural Networks
- Neural networks and rational functions
- Analysis and Design of Convolutional Networks via Hierarchical Tensor Decompositions
- Exploration of Neural Machine Translation in Autoformalization of Mathematics in Mizar
- Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization
- Efficient Representation of Low-Dimensional Manifolds using Deep Networks
- Is Deeper Better only when Shallow is Good?
- Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions
- Robust Training and Initialization of Deep Neural Networks: An Adaptive Basis Viewpoint
- Reverse-Engineering Deep ReLU Networks
- Representing smooth functions as compositions of near-identity functions with implications for deep network optimization
- Deep vs. shallow networks : An approximation theory perspective
- Provably Good Solutions to the Knapsack Problem via Neural Networks of Bounded Size
- The universal approximation power of finite-width deep ReLU networks
- Depth-Width Trade-offs for ReLU Networks via Sharkovsky's Theorem
- Survey of Expressivity in Deep Neural Networks
- Dissecting Deep Neural Networks
- The Nonlinearity Coefficient - Predicting Generalization in Deep Neural Networks
- Approximation power of random neural networks
- Implicit Rugosity Regularization via Data Augmentation
- CNNs Avoid Curse of Dimensionality by Learning on Patches
- Deep ReLU Networks Preserve Expected Length
- Depth-Width Trade-offs for Neural Networks via Topological Entropy
- Optimal Function Approximation with Relu Neural Networks
- Neural Networks with Small Weights and Depth-Separation Barriers
- Sharp bounds for the number of regions of maxout networks and vertices of Minkowski sums
- Hierarchically Compositional Tasks and Deep Convolutional Networks
- Approximating Lipschitz continuous functions with GroupSort neural networks
- Better Depth-Width Trade-offs for Neural Networks through the lens of Dynamical Systems
- Deep Learning Works in Practice. But Does it Work in Theory?
- Information-Theoretic Lower Bounds for Compressive Sensing with Generative Models
- When Hardness of Approximation Meets Hardness of Learning
- Universal Approximation of Residual Flows in Maximum Mean Discrepancy
- Collocation approximation by deep neural ReLU networks for parametric elliptic PDEs with lognormal inputs
- Limiting Network Size within Finite Bounds for Optimization
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- Depth-Adaptive Neural Networks from the Optimal Control viewpoint
- Learning Boolean Circuits with Neural Networks
- Efficient Deep Learning of GMMs
- Depth Enables Long-Term Memory for Recurrent Neural Networks
- Expressivity of Neural Networks via Chaotic Itineraries beyond Sharkovsky's Theorem
- A simple geometric proof for the benefit of depth in ReLU networks
- On Symmetry and Initialization for Neural Networks
- On the Expected Complexity of Maxout Networks