The power of deeper networks for expressing natural functions
arXiv:1705.05502
Abstract
It is well-known that neural networks are universal approximators, but that deeper networks tend in practice to be more powerful than shallower ones. We shed light on this by proving that the total number of neurons required to approximate natural classes of multivariate polynomials of variables grows only linearly with for deep neural networks, but grows exponentially when merely a single hidden layer is allowed. We also provide evidence that when the number of hidden layers is increased from to , the neuron requirement grows exponentially not with but with , suggesting that the minimum number of layers required for practical expressibility grows only logarithmically with .
Replaced to match version published at ICLR 2018. 14 pages, 2 figs
References in corpus (4)
Cited by in corpus (35)
- Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations
- A Review on Deep Learning in Medical Image Reconstruction
- Approximating Continuous Functions by ReLU Nets of Minimal Width
- Are All Layers Created Equal?
- The Static Local Field Correction of the Warm Dense Electron Gas: An ab Initio Path Integral Monte Carlo Study and Machine Learning Representation
- Nonlinear Approximation via Compositions
- Analyzing Upper Bounds on Mean Absolute Errors for Deep Neural Network Based Vector-to-Vector Regression
- ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction
- Improving Galaxy Clustering Measurements with Deep Learning: analysis of the DECaLS DR7 data
- On Interpretability of Artificial Neural Networks: A Survey
- A Selective Overview of Deep Learning
- D2RL: Deep Dense Architectures in Reinforcement Learning
- Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition
- Convergence Rates of Variational Inference in Sparse Deep Learning
- Universal Deep Beamformer for Variable Rate Ultrasound Imaging
- Deep Neural Networks for Choice Analysis: A Statistical Learning Theory Perspective
- Deep ReLU Networks Preserve Expected Length
- Deep Residual Mixture Models
- Deep Neural Networks for Choice Analysis: Extracting Complete Economic Information for Interpretation
- A Differential Topological View of Challenges in Learning with Feedforward Neural Networks
- Optimal Function Approximation with Relu Neural Networks
- Translating Diffusion, Wavelets, and Regularisation into Residual Networks
- On a Sparse Shortcut Topology of Artificial Neural Networks
- Tangent Space Separability in Feedforward Neural Networks
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Depth separation beyond radial functions
- A Machine-Learning Surrogate Model for ab initio Electronic Correlations at Extreme Conditions
- Realizing data features by deep nets
- On the approximation of functions by tanh neural networks
- Adaptive Variational Bayesian Inference for Sparse Deep Neural Network
- Fast suppression of classification error in variational quantum circuits
- MAGI-X: Manifold-Constrained Gaussian Process Inference for Unknown System Dynamics
- Expressive power of outer product manifolds on feed-forward neural networks
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks
- Interpretations of Deep Learning by Forests and Haar Wavelets