Learning Functions: When Is Deep Better Than Shallow
arXiv:1603.00988
Abstract
While the universal approximation property holds both for hierarchical and shallow networks, we prove that deep (hierarchical) networks can approximate the class of compositional functions with the same accuracy as shallow networks but with exponentially lower number of training parameters as well as VC-dimension. This theorem settles an old conjecture by Bengio on the role of depth in networks. We then define a general class of scalable, shift-invariant algorithms to show a simple and natural set of requirements that justify deep convolutional networks.
References in corpus (2)
Cited by in corpus (38)
- Social physics
- Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification
- A Review on Deep Learning in Medical Image Reconstruction
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Tensor Methods in Computer Vision and Deep Learning
- FETCH: A deep-learning based classifier for fast transient classification
- Wavefield reconstruction inversion via physics-informed neural networks
- The power of deeper networks for expressing natural functions
- Depth Creates No Bad Local Minima
- Extrapolation of nuclear structure observables with artificial neural networks
- Efficient Deep Learning Techniques for Multiphase Flow Simulation in Heterogeneous Porous Media
- BCR-Net: a neural network based on the nonstandard wavelet form
- A Selective Overview of Deep Learning
- Prediction of numerical homogenization using deep learning for the Richards equation
- A multiscale neural network based on hierarchical matrices
- Streaming Normalization: Towards Simpler and More Biologically-plausible Normalizations for Online and Recurrent Learning
- Analysis and Design of Convolutional Networks via Hierarchical Tensor Decompositions
- Butterfly-Net: Optimal Function Representation Based on Convolutional Neural Networks
- Representing smooth functions as compositions of near-identity functions with implications for deep network optimization
- A Framework for Searching for General Artificial Intelligence
- PolyGAN: High-Order Polynomial Generators
- Deep Learning and Hierarchal Generative Models
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- Prediction of Discretization of GMsFEM using Deep Learning
- The Representation Power of Neural Networks: Breaking the Curse of Dimensionality
- Deep Global Model Reduction Learning in Porous Media Flow Simulation
- On the Compressive Power of Deep Rectifier Networks for High Resolution Representation of Class Boundaries
- A multiscale neural network based on hierarchical nested bases
- Deep Learning: Generalization Requires Deep Compositional Feature Space Design
- From Deep to Shallow: Transformations of Deep Rectifier Networks
- Deep Neural Networks for Choice Analysis: Architectural Design with Alternative-Specific Utility Functions
- Limiting Network Size within Finite Bounds for Optimization
- On the space-time expressivity of ResNets
- Grow-Push-Prune: aligning deep discriminants for effective structural network compression
- General AI Challenge - Round One: Gradual Learning
- The staircase property: How hierarchical structure can guide deep learning
- Learning the mapping : the cost of finding the needle in a haystack
- Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning