Benefits of depth in neural networks
arXiv:1602.04485
Abstract
For any positive integer , there exist neural networks with layers, nodes per layer, and distinct parameters which can not be approximated by networks with layers unless they are exponentially large --- they must possess nodes. This result is proved here for a class of nodes termed "semi-algebraic gates" which includes the common choices of ReLU, maximum, indicator, and piecewise polynomial functions, therefore establishing benefits of depth against not just standard networks with ReLU gates, but also convolutional networks with ReLU and maximization gates, sum-product networks, and boosted decision trees (in this last case with a stronger separation: total tree nodes are required).
To appear, COLT 2016. For a simplified version, see http://arxiv.org/abs/1509.08101
References in corpus (1)
Cited by in corpus (49)
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Deep Residual Learning for Compressed Sensing CT Reconstruction via Persistent Homology Analysis
- Optimal approximation of continuous functions by very deep ReLU networks
- Understanding Deep Neural Networks with Rectified Linear Units
- Review: Deep Learning in Electron Microscopy
- The Shattered Gradients Problem: If resnets are the answer, then what is the question?
- Exploiting deep residual networks for human action recognition from skeletal data
- Two-hidden-layer Feedforward Neural Networks are Universal Approximators: A Constructive Approach
- DeLighT: Deep and Light-weight Transformer
- Effect of Depth and Width on Local Minima in Deep Learning
- On the ability of neural nets to express distributions
- Learning to Recognize 3D Human Action from A New Skeleton-based Representation Using Deep Convolutional Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- DeepOPF: A Deep Neural Network Approach for Security-Constrained DC Optimal Power Flow
- Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks
- On the Number of Linear Regions of Convolutional Neural Networks
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Understanding and Resolving Performance Degradation in Graph Convolutional Networks
- Transport Analysis of Infinitely Deep Neural Network
- Convolutional Rectifier Networks as Generalized Tensor Decompositions
- Deep artifact learning for compressed sensing and parallel MRI
- Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank
- Is Deeper Better only when Shallow is Good?
- Butterfly-Net: Optimal Function Representation Based on Convolutional Neural Networks
- Exponential Convergence of the Deep Neural Network Approximation for Analytic Functions
- Deep Graph Neural Networks with Shallow Subgraph Samplers
- Stability and Generalization of Graph Convolutional Neural Networks
- Towards Understanding Hierarchical Learning: Benefits of Neural Representations
- Deep Learning and Hierarchal Generative Models
- Optimal Nonparametric Inference via Deep Neural Network
- Partition of unity networks: deep hp-approximation
- Neural Networks with Small Weights and Depth-Separation Barriers
- Deep Online Convex Optimization with Gated Games
- Function approximation by deep networks
- Size and Depth Separation in Approximating Benign Functions with Neural Networks
- The Connection Between Approximation, Depth Separation and Learnability in Neural Networks
- Beyond Deep Residual Learning for Image Restoration: Persistent Homology-Guided Manifold Simplification
- Depth separation beyond radial functions
- On the Universal Approximation Property and Equivalence of Stochastic Computing-based Neural Networks and Binary Neural Networks
- When Hardness of Approximation Meets Hardness of Learning
- Deep Learning Works in Practice. But Does it Work in Theory?
- Construction of neural networks for realization of localized deep learning
- Linearly Constrained Weights: Reducing Activation Shift for Faster Training of Neural Networks
- Learning Boolean Circuits with Neural Networks
- Approximation smooth and sparse functions by deep neural networks without saturation
- The emergent algebraic structure of RNNs and embeddings in NLP
- A simple geometric proof for the benefit of depth in ReLU networks
- Generalization and Expressivity for Deep Nets
- Normalization effects on shallow neural networks and related asymptotic expansions