The Power of Depth for Feedforward Neural Networks
arXiv:1512.03965
Abstract
We show that there is a simple (approximately radial) function on , expressible by a small 3-layer feedforward neural networks, which cannot be approximated by any 2-layer network, to more than a certain constant accuracy, unless its width is exponential in the dimension. The result holds for virtually all known activation functions, including rectified linear units, sigmoids and thresholds, and formally demonstrates that depth -- even if increased by 1 -- can be exponentially more valuable than width for standard feedforward neural networks. Moreover, compared to related results in the context of Boolean functions, our result requires fewer assumptions, and the proof techniques and construction are very different.
Accepted to COLT 2016; Fixed a bug in the proof of claim 2 (now requiring the mild assumption that the activations are polynomially bounded); Other minor revisions
References in corpus (2)
Cited by in corpus (50)
- PhysNet: A Neural Network for Predicting Energies, Forces, Dipole Moments and Partial Charges
- Data Driven Governing Equations Approximation Using Deep Neural Networks
- Exponential expressivity in deep neural networks through transient chaos
- Provable approximation properties for deep neural networks
- FETCH: A deep-learning based classifier for fast transient classification
- Convergence of the Deep BSDE Method for Coupled FBSDEs
- The Modern Mathematics of Deep Learning
- On the Expressive Power of Deep Learning: A Tensor Analysis
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
- Review: Deep Learning in Electron Microscopy
- Are All Layers Created Equal?
- Benefits of depth in neural networks
- On the Origin of Deep Learning
- On the Expressive Power of Deep Neural Networks
- Convolutional Recurrent Neural Networks for Music Classification
- Analyzing Upper Bounds on Mean Absolute Errors for Deep Neural Network Based Vector-to-Vector Regression
- Semi-supervised Feature Learning For Improving Writer Identification
- On the ability of neural nets to express distributions
- A simple and efficient architecture for trainable activation functions
- Inductive Bias of Deep Convolutional Networks through Pooling Geometry
- Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
- Toward a Better Monitoring Statistic for Profile Monitoring via Variational Autoencoders
- CTU Depth Decision Algorithms for HEVC: A Survey
- Explaining Deep Neural Networks
- Analysis and Design of Convolutional Networks via Hierarchical Tensor Decompositions
- Convolutional Rectifier Networks as Generalized Tensor Decompositions
- Microparticle cloud imaging and tracking for data-driven plasma science
- Efficient Representation of Low-Dimensional Manifolds using Deep Networks
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Towards Understanding Theoretical Advantages of Complex-Reaction Networks
- Deep vs. shallow networks : An approximation theory perspective
- Deep Nonparametric Regression on Approximate Manifolds: Non-Asymptotic Error Bounds with Polynomial Prefactors
- Survey of Expressivity in Deep Neural Networks
- CovNet: Covariance Networks for Functional Data on Multidimensional Domains
- Deep Stochastic Configuration Networks with Universal Approximation Property
- CNNs Avoid Curse of Dimensionality by Learning on Patches
- Approximation power of random neural networks
- The Nonlinearity Coefficient - Predicting Generalization in Deep Neural Networks
- The Role of Information Complexity and Randomization in Representation Learning
- Representational Power of ReLU Networks and Polynomial Kernels: Beyond Worst-Case Analysis
- From Shallow to Deep Interactions Between Knowledge Representation, Reasoning and Machine Learning (Kay R. Amel group)
- Deep Online Convex Optimization with Gated Games
- Physics-regularized neural network of the ideal-MHD solution operator in Wendelstein 7-X configurations
- Second Harmonic Imaging Enhanced by Deep Learning Decipher
- On the Compressive Power of Deep Rectifier Networks for High Resolution Representation of Class Boundaries
- Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
- Limiting Network Size within Finite Bounds for Optimization
- Generalization and Expressivity for Deep Nets
- Exploiting Nontrivial Connectivity for Automatic Speech Recognition