Approximating Continuous Functions by ReLU Nets of Minimal Width
arXiv:1710.11278
Abstract
This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed what is the minimal width so that neural nets with ReLU activations, input dimension , hidden layer widths at most and arbitrary depth can approximate any continuous, real-valued function of variables arbitrarily well? It turns out that this minimal width is exactly equal to That is, if all the hidden layer widths are bounded by , then even in the infinite depth limit, ReLU nets can only express a very limited class of functions, and, on the other hand, any continuous function on the -dimensional unit cube can be approximated to arbitrary precision by ReLU nets in which all hidden layers have width exactly Our construction in fact shows that any continuous function can be approximated by a net of width . We obtain quantitative depth estimates for such an approximation in terms of the modulus of continuity of .
v2. 13p. Extended main result to higher dimensional output. Comments welcome
References in corpus (3)
Cited by in corpus (53)
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- A Review on Deep Learning in Medical Image Reconstruction
- ResNet with one-neuron hidden layers is a Universal Approximator
- Optimal approximation of continuous functions by very deep ReLU networks
- The Modern Mathematics of Deep Learning
- Review: Deep Learning in Electron Microscopy
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- Nonlinear Approximation and (Deep) ReLU Networks
- Nonlinear Approximation via Compositions
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Two-hidden-layer Feedforward Neural Networks are Universal Approximators: A Constructive Approach
- Universal approximations of permutation invariant/equivariant functions by deep neural networks
- Are Transformers universal approximators of sequence-to-sequence functions?
- The Calabi-Yau Landscape: from Geometry, to Physics, to Machine-Learning
- Neural Ordinary Differential Equation Control of Dynamics on Graphs
- Densely connected neural networks for nonlinear regression
- DeepOPF: A Deep Neural Network Approach for Security-Constrained DC Optimal Power Flow
- Universal Approximation with Deep Narrow Networks
- On the Number of Linear Regions of Convolutional Neural Networks
- A Shallow Ritz Method for Elliptic Problems with Singular Sources
- Robust Training and Initialization of Deep Neural Networks: An Adaptive Basis Viewpoint
- How hard is to distinguish graphs with graph neural networks?
- On the Modularity of Hypernetworks
- Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions
- A Corrective View of Neural Networks: Representation, Memorization and Learning
- Approximation in shift-invariant spaces with deep ReLU neural networks
- Provably Good Solutions to the Knapsack Problem via Neural Networks of Bounded Size
- Predicting the transverse emittance of space charge dominated beams using the phase advance scan technique and a fully connected neural network
- Minimum Width for Universal Approximation
- Provable Memorization via Deep Neural Networks using Sub-linear Parameters
- Implicit Rugosity Regularization via Data Augmentation
- Approximation capabilities of neural networks on unbounded domains
- On Universal Equivariant Set Networks
- COLD: Concurrent Loads Disaggregator for Non-Intrusive Load Monitoring
- Sharp Representation Theorems for ReLU Networks with Precise Dependence on Depth
- Model Complexity of Deep Learning: A Survey
- Arbitrary-Depth Universal Approximation Theorems for Operator Neural Networks
- Optimal Function Approximation with Relu Neural Networks
- Universal Approximation of Input-Output Maps by Temporal Convolutional Nets
- A Topological Framework for Deep Learning
- Integration of Fractional Order Black-Scholes Merton with Neural Network
- Error Estimation and Correction from within Neural Network Differential Equation Solvers
- Deep Learning Framework for Hybrid Analog-Digital Signal Processing in mmWave Massive-MIMO Systems
- Quantifying Epistemic Uncertainty in Deep Learning
- Mathematical Foundations of Regression Methods for the approximation of the Forward Initial Margin
- Meta Internal Learning
- Algebraically-Informed Deep Networks (AIDN): A Deep Learning Approach to Represent Algebraic Structures
- Topological Deep Learning: Classification Neural Networks
- Fourier Neural Networks for Function Approximation
- When Can Neural Networks Learn Connected Decision Regions?
- Universality of Gradient Descent Neural Network Training
- Prevention is Better than Cure: Handling Basis Collapse and Transparency in Dense Networks
- Error estimate for a universal function approximator of ReLU network with a local connection